The margin column holds each section's figure, cell and caveat, as it does today. In the few sections where motion says more than a still, a stage appears at the top of the margin, stays put while you scroll that section, and steps with the paragraph you are reading. The example is a distillation paper: the temperature, the loss and the layer matching are animated; the rest stay figures.
A large model, or an ensemble of them, learns a function that a much smaller model could represent but would not find on its own. The paper's claim is that the knowledge sits in the big model's output distribution, not in its weights, and that a small model can be trained to copy that distribution.¶1
The method is two models and one extra loss. A teacher, already trained, and a student, small enough to deploy, see the same inputs. The student is trained to match the teacher's softened predictions as well as the true labels.
It works because wrong answers carry information. A teacher that puts 0.9 on cat and 0.08 on dog has told the student that this cat looks like a dog, something a one-hot label never says. Who should care: anyone shipping a model to a budget the training run never had.
Training and serving want different models. In training you are free to be extravagant: train many models, average them, use far more parameters than the data strictly needs, because extra capacity makes the optimisation easier. Watch a forward pass through such a model on the right: every layer touches every unit, and the cost is the number of arrows.
The extreme case is an ensemble. Ten models, ten forward passes, and the answer is the average. In 2015 this was the reliable way to win a benchmark, and it was ten times the cost of one model at serving time.
The serving budget is fixed by latency and memory on a phone or a datacentre at scale. The question the paper asks is whether the function the ensemble has learned, as opposed to its parameters, can be moved into one small model.
Start with how a classifier is usually trained. The target for one image is a hard label: 1 on the right class, 0 everywhere else. The stage shows that for a picture of a cat over five classes.
Now look at what a trained teacher actually outputs for the same image. Most of the mass is on cat, but a little goes to dog, far less to car and truck, and almost nothing to ship. That ordering is not noise. It is the teacher's similarity structure over classes, learned from millions of images.
The paper calls this dark knowledge: the information in the small probabilities. One soft target constrains the student far more than one hard label, because it says something about every class at once. The trouble is that at the usual softmax these small probabilities are tiny, around 10⁻⁶, and contribute almost nothing to a cross-entropy gradient.
The fix is to divide the logits by a temperature T before the softmax, in both the teacher and the student:
At T = 1 this is the ordinary softmax and the distribution is peaked. Raise T and the logits are squeezed together, so the distribution flattens and the small probabilities become visible to the loss. The stage raises T as you read on; watch the dog and car bars grow while the ordering stays the same.
At very high T the softmax becomes nearly uniform and the information is gone again, so there is a useful band. The paper used T between 2 and 20, with smaller students preferring lower T because they cannot fit all of the teacher's structure anyway.
Because dividing by T scales gradients by 1/T², the paper multiplies the soft loss by T² so that the two terms stay comparable when T is tuned. That is a bookkeeping detail, but one that matters when you implement it.
Both networks see the same batch. The teacher runs forward once with its weights frozen; the student runs forward and will be updated. On the stage, the same input enters both stacks and each produces logits at the top.
The first term compares the two softened outputs. Teacher logits at temperature T give a soft distribution; student logits at the same T give another; the loss is the cross-entropy between them, which is the KL divergence up to a constant. Only the final layer's outputs are compared. Nothing inside either network is looked at.
The second term is the ordinary cross-entropy against the hard label, at T = 1. It keeps the student honest when the teacher is wrong and gives it a sharp decision at serving time. The paper found a small weight on this term works best.
Gradients flow back through the student only. The teacher is a fixed function during distillation, so it is drawn greyed out once the backward pass starts. This is why distillation is cheap to add: one extra forward pass through a model you already have.
This paper matches outputs only, but the obvious next question is whether the student should also match what the teacher computes on the way. Later work, starting with FitNets in 2015, does exactly that, and it is the version most people mean today when they say they distilled an intermediate layer. The stage shows a six-layer teacher beside a three-layer student.
The teacher's second block is paired with the student's first. The two feature maps have different widths, 512 against 256 here, so a small learned regressor projects the student's activations into the teacher's space before a mean-squared-error loss is taken between them.
The fourth teacher block is paired with the student's second block in the same way. Which pairs to choose is a design decision; the usual rule is to spread them evenly through the depth so that the student is guided throughout, rather than only told the answer.
The final pair is the two output layers, which is this paper's loss again. Putting it all together, the student is pulled toward the teacher at three points in its depth, each with its own weight in the total loss.
On MNIST, a large teacher with two hidden layers of 1200 units makes 67 test errors. A small student of two layers of 800 units trained on labels alone makes 146. Watch the student's error curve on the stage as it trains on hard labels only.
The same small student, trained with soft targets from the teacher at T = 20, makes 74 errors. The second curve shows that run. Almost the whole gap to the teacher closes, with no change to the student's architecture.
The striking experiment removes every example of the digit 3 from the student's training set. Distilled from the teacher, the student still recognises 98.6% of the 3s it has never seen, because the teacher's soft targets on other digits say what a 3 is not.
for soft in (False, True):
errs = train_student(soft_targets=soft, T=20)
print("soft" if soft else "hard", errs[-1])
Output distillation is standard practice and the recipe has barely changed. Intermediate-layer matching, attention-map matching and relational losses were added by later work and are usually combined with it.
For language models the method moved to sequence level: DistilBERT and TinyBERT in 2019 distilled encoder models with layer matching, and current instruction-tuned small models are routinely trained on a larger model's generations, which is distillation through samples rather than through logits.
What a reader should use today: the T² soft-target loss as written, a modest weight on the hard label, and one or two intermediate matches if the student is much shallower than the teacher.
What the mock above now does, chosen from the options below it.
Every section keeps its figure in the margin, drawn as today. Only some sections get a scene: the ones Claude marks as worth animating when it writes the page, a mechanism or a loss rather than a table or a timeline, and the ones the reader asks for. The page is never without its stills.
The stage appears at the top of the margin when you enter a section with a scene, follows its paragraphs, and leaves when the section ends. In every other section the margin is exactly what it is now.
A figure Claude thinks worth animating carries Animate · Claude suggests it here; the outline marks the section with a hollow play. Pressing it, or asking in the bar, streams the scene in node by node and saves it with the page. Any other figure can be asked for the same way; Claude just does not offer.
The figure is the scene at rest. On the stage, the Still tab is its resting frame; steps, bound to paragraphs, animate it. The figure card stays in the margin under the stage, marked as the source. Keep and the notes always get the still.
A section with several pictures gets thumbnails under the stage: scene, still, the paper's own figure. The current one is outlined; a click shows any of them.
A scene can name a cell. Until it runs, the scene shows Claude's illustrative values; after, the real ones, and the header says which cell and machine. The training section shows this.
The stage's menu renders the scene to WebM or GIF in the browser and keeps a frame as a still. Nothing is generated as a video file; a video is made from the scene when someone wants one.
What the stage is, where it could sit, what moves it, and what Claude would have to write for it to exist.
A section of the explanation gets, besides its figure and its cell, a scene: a small animation with named steps. The stage shows the scene of the section you are in and the step of the paragraph you are on. It stays on screen while the prose scrolls, so the text and the picture are always read together.
The Margin layout already has the column. The stage takes its top and stays; the section's static figures and cells keep scrolling beneath it.
Figure, cell and caveat are tabs on one sticky stage, as in the mock's header. Nothing scrolls in the margin at all; the stage swaps what it shows as you read.
A draggable card over the bottom corner, the way the reader's floating windows already work. Works in every layout, including Notebook and Beside the paper, where there is no margin.
On the phone layout the stage becomes a short sticky band above the prose, landscape, with the caption to the right of the picture. The mock does this below 1100px.
Each paragraph is a step. The paragraph crossing the reading line (a third of the way down) picks the step; the scene animates to that state and idles there. Reading backwards steps backwards.
Scroll position within the section maps to time within the scene, like an Apple product page. The slider in the mock is this, by hand.
Entering a section plays its scene once, then it loops slowly. The paragraph steps are ignored.
The page already draws four fenced blocks specially. A fifth, motion, would be a declarative scene rather than code, so it can be validated, themed and redrawn without running anything untrusted. Nodes, groups, edges, and per-step properties; the page owns the layout and the tweening.
```motion title="The distillation loss" steps="forward,soft match,hard label,backward"
nodes:
x: {kind: input, label: "x"}
teacher: {kind: stack, layers: 6, label: "teacher", frozen: true}
student: {kind: stack, layers: 3, label: "student"}
zT: {kind: dist, label: "softmax(z_T / T)", from: teacher}
zS: {kind: dist, label: "softmax(z_S / T)", from: student}
y: {kind: label, label: "y"}
edges:
- {from: x, to: [teacher, student], flow: forward}
- {from: zT, to: zS, kind: compare, label: "KL", step: 1}
- {from: y, to: zS, kind: compare, label: "CE", step: 2}
- {from: zS, to: student, flow: backward, step: 3}
steps:
0: {caption: "The same batch runs through both networks."}
1: {caption: "Only the two softened outputs are compared.", highlight: [zT, zS]}
2: {caption: "A small hard-label term keeps the student honest.", highlight: [y, zS]}
3: {caption: "Gradients reach the student only.", dim: [teacher]}
```
Steps are bound to paragraphs in order, or by name with a marker in the prose (the ¶1 in the mock's first paragraph). A vocabulary of a dozen node kinds covers most ML papers: stack, layer, tensor, dist, curve, grid, timeline, token row, attention matrix. Papers outside ML get fewer scenes and that is fine; a section without a motion block shows its figure on the stage instead.
| Kind of section | Scene that fits |
|---|---|
| Architecture | Flow through a stack: packets along edges, layers lighting in order, a tensor's shape changing at each. |
| Loss or objective | Two outputs and the comparison between them, then the gradient path drawn back. Which layers meet which. |
| A hyperparameter | One slider, one live chart: T and the softmax, the learning rate and a curve, the context window and a mask. |
| Training | Curves drawn over time, the two runs the text contrasts. |
| Attention, routing, mixing | A matrix or bipartite graph whose weights fill in as the step moves. |
| Since then | A timeline, each later paper a dot the caption names. |
A still and a scene do different jobs. A still shows everything at once and is what you keep, print and come back to. A scene shows one thing at a time and is only useful while you read. The page should have both, and the question is only where each goes and whether Claude draws them twice.
The margin keeps its rows. Each still, snip, caveat and cell sits beside its paragraph, pushed down by the stage's height, and comes out from under the stage as you read on.
The figure Claude already draws is the scene at rest. A motion block names elements in that SVG by id and says what each step does to them: light this, move that, draw this edge. No step, and it is a still as today.
Figures and snips go into the reading column at full width, the way Notebook lays them out, and the margin holds only the stage, cells and caveats. What moves stays put; what is still scrolls with the text.
Both sticky. The lower pane crossfades to the still that belongs to the paragraph being read, so a still is never behind the stage and never trails.
Thumbnails of everything the section has: the scene, its stills, the paper's own figures. The one that belongs to the paragraph is outlined as you read; a click shows any of them on the stage and pins it.
Above about 1500px the margin splits: stills beside their paragraphs in a narrow column, the stage sticky in a third. Below that width it collapses to the first option.
The stage in the mock has a Paper tab. For any section Claude can point at a figure in the paper, a snip from the PDF is a tab beside Claude's redraw and the scene, with the page number and a link into the paper. Three pictures of the same thing, from the most trusted to the most explanatory, and the reader can check one against the other. The reader already snips PDFs, so this is the existing snip, placed by Claude.
| Case | What the reader sees |
|---|---|
| Section with one figure | The figure is the scene at rest; steps animate it; Figure tab shows it still. Nothing in the margin but cells and caveats. |
| Section with several figures | The main one is the scene; the others are stills beside their paragraphs beneath the stage, and thumbnails on the filmstrip. |
| The paper's own figure | Paper tab on the stage, and a snip in the margin where Claude cites it. |
| No motion block | The stage shows the section's figure, still. A section with no figure folds the stage to a one-line strip. |
| Keep, notes, Markdown export | Always a still: the resting frame or the current step as SVG, with its caption. Motion never leaves the page. |
| Phone | A short strip with the still, tap to play; the margin items go inline under their paragraphs. |
Recommendation. One drawing, two states, as the authoring model: it removes the duplication at the root and makes every scene also a figure. Stage on top with the margin flowing under it as the default placement, with the stage kept short and the filmstrip for sections that have more than one picture. The Paper tab for figures Claude cites. Stills in the prose column as a separate "Stage" layout for readers who prefer it.
Hovering a highlighted term in the prose (teacher, temperature, regressor) pulses the node it names on the stage. The reverse too: hovering a node underlines its term.
Hold a scene while reading on. The live dot in the stage header goes grey and the text says pinned to §05. Scrolling into a new section shows a small follow nudge rather than switching.
Puts the scene's spec and current step into the Ask bar as scope, the way a cell's output does now: why is the teacher greyed out here?
Asks Claude for a different scene for this section: show it with real feature map sizes, animate the backward pass slower. The block is rewritten like any other section edit.
Keeps the scene's current frame as a still, with its caption and step, in your notes. The notes clip renders it as a figure.
The ⤢ button opens the scene in a floating window with the caption under it, for a projector or a second look.
Margin vs stage. The paragraph-adjacent figure is the Tufte idea the Margin layout is built on. The stage breaks adjacency for whatever used to sit at the top. The mock's answer is to keep both, stage above, figures below, but it is worth deciding whether a section's figure simply becomes the last step of its scene.
How much motion. Idling animation beside prose pulls the eye. The mock idles gently (packets moving, a curve breathing) and has a pause. A stricter version animates only on a step change and is otherwise still.
Where the scenes come from. Claude writes them with the explanation, which makes a page slower and longer to write. Alternatively they are written on demand the first time a section is read, with a drawing… state on the stage, which keeps the first explanation as fast as it is now.
Scope per paper. Not every section needs a scene. Two to four per paper, on the method and the loss, is probably right; the outline's small play icon marks which have one.