Scaled Dot-Product Attention Explorer Interactive
Adjust sequence length, key dimension and temperature to see how the attention weight matrix changes. Hover cells to inspect probabilities. Random embeddings are generated with a reproducible seed.
Parameters
τ: user temperature scaling (additional to the canonical \(\sqrt{d_k}\) factor)
"Freeze One Token" keeps token 0 fixed and regenerates others—useful to observe stability.
Attention Weights Heatmap
Value Output (Optional View)
{{valueVectors | json}}
Explanation of Core Concepts & Controls
This explorer lets you see how the scaled dot‑product attention weights react when you change sequence length, the key/query dimensionality dk, and a user‑defined temperature τ. Each row of the matrix is a softmax distribution over keys for one query token.
1. Scaled Dot-Product Attention
Given query matrix Q ∈ ℝ^{L×d_k}, key matrix K ∈ ℝ^{L×d_k}, value matrix V ∈ ℝ^{L×d_v} (here d_v = d_k / 2 for display compactness), raw scores are S = QKT. We then scale by √d_k (to keep variance of dot products roughly constant as d_k grows) and additionally divide by user temperature τ. So the logits are S / (√d_k · τ). Applying a row‑wise softmax produces attention weights A; each row of A sums to 1.
2. Sequence Length (L)
Increasing L expands both matrix dimensions: more query positions (rows) and more key positions (columns). The amount of probability mass per row is fixed (always sums to 1) so as you lengthen the sequence, typical individual weights become smaller unless there is a strong alignment. Computationally, attention cost is O(L² d_k) for forming the score matrix plus softmax.
3. Key / Query Dimension (dk)
This sets the length of each embedding vector in Q and K. Larger d_k allows more representational capacity but also increases raw dot‑product magnitude variance. The canonical scaling factor 1/√d_k normalizes those magnitudes so that softmax does not become saturated (which would yield near one‑hot rows and vanishing gradients). Try pushing d_k high and toggling the temperature to see how scaling preserves reasonable distributions.
4. Temperature (τ)
Lower τ < 1 effectively multiplies logits, sharpening the softmax (peaky attention). Higher τ > 1 divides logits, smoothing distributions and increasing entropy. This is analogous to sampling temperature in decoding, but applied earlier—to attention compatibility scores.
5. Freeze One Token
Regenerates all tokens except token 0. This lets you inspect how a single fixed query/key vector’s attention row changes when its context (all the other tokens) changes. Useful for intuition about relative vs absolute embedding geometry.
6. Export JSON
Downloads the current attention matrix plus parameter metadata so you can analyze it offline (e.g., compute entropy per row, argmax structure, or simulate multi‑head variants).
7. Value Projection
After computing weights, each output vector is Σj A[i,j] · V[j]. We display rounded numbers so you can glimpse how smoother vs sharper rows lead to blended vs near‑copied value vectors.
8. Practical Notes
- Real Transformers use multi‑head attention: separate learned linear projections form multiple (Q,K,V) triplets; outputs are concatenated.
- Scaling by
√d_kis critical; without it, larged_kcollapses softmax into extreme spikes. - Entropy of each row can diagnose whether attention is over‑confident; you could extend this page to compute it.
- Temperature here is an experimental extra knob (not standard inside attention) to illustrate logit sensitivity.
- Reproducible pseudo‑random numbers: a simple linear congruential generator ensures identical results for same parameters until you re‑randomize.
Return to the main formula index to explore Cross‑Entropy and Adam pages.