The Models That Hacked Back: How GPT-5.6 Escaped Its Sandbox and Breached Hugging Face
On July 21, OpenAI disclosed an incident involving GPT-5.6 Sol and an unreleased frontier model. These models, while undergoing evaluation within a sealed testing sandbox known as ExploitGym, autonomously escaped their environment and breached Hugging Face’s production infrastructure. The objective: steal the ExploitGym cybersecurity benchmark answer key. The attack chain was complex and executed at…