Cross-Entropy Explorer (Text-Aware) index463
We derive an empirical distribution from your text in index461. The true class at each position is the actual observed word. We simulate a model that assigns an adjustable logit advantage Δ to each true word and temperature scale τ to the rest. Observe how average loss approaches entropy H(p) as advantage grows.
Editable Text & Parameters
Words: {{meta.total}}
Unique: {{meta.vocab}}
H(p): {{meta.entropy | number:3}}
PPL: {{meta.perplexity | number:2}}
Parameters
\(
\begin{aligned}
ext{Sequence Loss } \mathcal{L} &= - \sum_{t=1}^{L} \log p(y_t) \\
H(p) &= - \sum_{w \in V} p(w) \log p(w) \\
ext{Perplexity} &= e^{H(p)}
\end{aligned}
\)
Empirical H(p): {{meta.entropy | number:3}}
Perplexity: {{meta.perplexity | number:2}}
Uniform log|V|: {{meta.uniformCe | number:3}}
Avg Loss: {{metrics.meanLoss | number:4}}
Median: {{metrics.medianLoss | number:4}}
Min/Max: {{metrics.minLoss | number:3}} / {{metrics.maxLoss | number:3}}
As Δ → large, the model concentrates probability on the true class → loss approaches 0. With Δ=0 and τ=1, the model nears (noisy) uniform → loss ≈ log|V|.
Distribution for Current Position
Top-k classes; remaining grouped as Other.
Class {{hover.idx}}: p={{hover.p|number:4}} (TRUE)
Per-Token Loss Table
| t | word | p_true | −log p |
|---|---|---|---|
| {{$index}} | {{row.word}} | {{row.pTrue|number:4}} | {{row.loss|number:4}} |