Cross-Entropy / NLL Explorer Interactive
Experiment with vocabulary size, sequence length, logit advantage, and temperature to see how cross-entropy loss changes.
Parameters
\(
\begin{aligned}
ext{Sequence Loss } \mathcal{L} &= - \sum_{t=1}^{L} \log p(y_t \mid x_{
Guidance: Increase Δ (advantage) to push more probability mass onto the correct class → lower loss.
Probability Distribution (Focused Token)
Showing top {{params.topK}} classes (remaining grouped as "Other"). Hover bars for details.
Class {{hover.index}}: p = {{hover.p|number:4}} (TRUE)
Per-Token Loss
| t | true cls | p_true | −log(p_true) |
|---|---|---|---|
| {{$index}} | {{row.truth}} | {{row.pTrue | number:4}} | {{row.loss | number:4}} |
Average Loss: {{metrics.meanLoss | number:4}}
Median Loss: {{metrics.medianLoss | number:4}}
Min/Max Loss: {{metrics.minLoss | number:3}} / {{metrics.maxLoss | number:3}}
Avg Entropy (focus dist): {{focusData.entropy | number:3}}
Loss vs Advantage Δ (Snapshot Curve)
Click "Capture Point" to append current (Δ, mean loss).