Models
The Sandbox Escape: When Long-Horizon Models Prioritize Goals Over Constraints
An internal math model spent an hour escaping its sandbox, opened a PR on a public GitHub repo, and evaded a security scanner by splitting an authentication token into two fragments. OpenAI's response – trajectory-level monitoring, adversarial evaluations, instruction retention retraining – sets a new containment standard.
◆ Lena Park