The alternatives considered: ← the pictures, with the design as built
Explained by Claude Opus 5 ExplanationImplementation MarginNotebookBeside the paper ▶ StageColab · no runtimeRuntimeMetricsRewrite ▾✕ Skip
✎ Ask Claude Opus 5 anything about this explanation, or how to change it… ( / )

03Scaled dot-product attention

Each token is turned into three vectors: a query (what am I looking for?), a key (what do I contain?) and a value (what do I hand over if chosen?). A token's output is an average of all the values, weighted by how well its query matches each key:

Attention(Q, K, V) = softmax(Q Kᵀ / √d_k) · V

That is the whole mechanism. The softmax turns match scores into weights that sum to one, so the output is a soft lookup: like a dictionary, but it returns a blend of the entries whose keys are closest.

The dot product of two random d-dimensional vectors has variance d. With d = 512 the scores are huge, the softmax puts almost all its weight on one key, and its gradient all but vanishes. Dividing by √d brings the variance back to 1 whatever the width.

⤡ Back to the margin Q K MatMul QKᵀ ÷ √d_k softmax × V V out weights sum to 1
The pipeline, left to right: match queries against keys, scale, normalise into weights, and use the weights to mix the values.
In [1] Self-attention over four tokens, from scratch ▶ Run in Colab
def attention(Q, K, V):
    scores = Q @ K.T / np.sqrt(K.shape[-1])
    w = np.exp(scores - scores.max(-1, keepdims=True))
    w /= w.sum(-1, keepdims=True)
    return w @ V, w
Figure 2 of 3 · Scaled dot-product attention Add to notes✕
Fit2×4×
The pipeline, left to right: match queries against keys, scale, normalise into weights, and use the weights to mix the values.
Close-up 2 of 3 + − zoom · drag to pan ⇥ next Esc back to the page