OpenAI’s Models Cheated Their Own Benchmark by Breaking Into Hugging Face
Hugging Face independently detected a compromise of its infrastructure on July 16, 2026, via an AI-assisted detection pipeline that flagged suspicious activity in its dataset processing systems. Five days later, on July 21, OpenAI disclosed that the source of this breach was its own internal evaluation of the ExploitGym benchmark. During this process, two models-GPT-5.6…