Three Core Mathematical Formulae Behind Modern LLMs
Interactive pages: index255.html, index254.html, and index253.html.
-
Scaled Dot-Product Attention Architecture
\\( \\text{Attention}(Q,K,V) = \\operatorname{softmax}\\left( \\frac{QK^{T}}{\\sqrt{d_k}} \\right) V \\)Captures contextual relationships between token representations by weighting values using query–key similarity.Open interactive: index452.html
-
Cross-Entropy (Negative Log-Likelihood) Loss Training Objective
\\( \\mathcal{L} = - \\sum_{t=1}^{T} \\log p(x_t \\mid x_{Measures how well the model predicts the next token; lower values indicate better predictive distribution alignment.Open interactive: index453.html
- Adam Optimizer Update Rule Optimization
\\[ \\begin{aligned} m_t &= \\beta_1 m_{t-1} + (1-\\beta_1) g_t \\\\ v_t &= \\beta_2 v_{t-1} + (1-\\beta_2) g_t^{2} \\\\ \\hat{m}_t &= \\frac{m_t}{1-\\beta_1^{t}}, \\quad \\hat{v}_t = \\frac{v_t}{1-\\beta_2^{t}} \\\\ \\theta_{t+1} &= \\theta_t - \\eta \\; \\frac{\\hat{m}_t}{\\sqrt{\\hat{v}_t}+\\epsilon} \\end{aligned} \\]Adaptive first & second moment estimation for stable, scale-invariant parameter updates.Open interactive: index454.htmlOverview
index452.html: Attention heatmap & parameter exploration.index453.html: Cross-entropy probability & loss visualization.index454.html: Adam vs SGD trajectory & loss comparison.
- Adam Optimizer Update Rule Optimization