03Scaled dot-product attention
Each token is turned into three vectors: a query (what am I looking for?), a key (what do I contain?) and a value (what do I hand over if chosen?). A token's output is an average of all the values, weighted by how well its query matches each key:
Attention(Q, K, V) = softmax(Q Kᵀ / √d_k) · V
That is the whole mechanism. The softmax turns match scores into weights that sum to one, so the output is a soft lookup: like a dictionary, but it returns a blend of the entries whose keys are closest.
The dot product of two random d-dimensional vectors has variance d. With d = 512 the scores are huge, the softmax puts almost all its weight on one key, and its gradient all but vanishes. Dividing by √d brings the variance back to 1 whatever the width.
In [1] Self-attention over four tokens, from scratch ▶ Run in Colab
def attention(Q, K, V):
scores = Q @ K.T / np.sqrt(K.shape[-1])
w = np.exp(scores - scores.max(-1, keepdims=True))
w /= w.sum(-1, keepdims=True)
return w @ V, w