LLM Core Formulae – Text Driven View
Enter or paste up to ~500 words of text. We derive sequence length, vocabulary size, and a frequency distribution to ground three fundamental formulae used in large language models: Scaled Dot‑Product Attention, Cross‑Entropy Loss, and the Adam Optimizer. The other pages (index462–index464) will read this shared text and tailor their interactive visualizations.
Shared Editable Text (Canonical Sample)
Formulae (Contextualized)
Attention(Q,K,V) = softmax( (Q KT) / ( √dk ) ) V
Sequence length L comes directly from the number of tokens extracted from your text. Larger L ⇒ larger attention matrix (L×L) and quadratic memory/time.
CE(y, p) = − Σt=1..L log p(yt | context)
If we approximate the empirical distribution of words in your text by
p(w), its entropy H(p) ≈ expected per‑token cross‑entropy for an optimal model. Perplexity = e^{H(p)}.Adam: θt+1=θt−α m̂t/(√(v̂t)+ε)
Vocabulary size influences embedding parameter count (≈|V|·d). Larger parameter spaces benefit from adaptive learning rates (Adam) for faster convergence than vanilla SGD.