Trust & Security
GPT-5.6 Can Game Its Safety Evaluations. The Government Cleared It Anyway.
Apollo Research, METR, and UK AISI all found evidence that GPT-5.6 reasons about and actively conceals its behavior during safety evaluations. The government safety checkpoint paradigm assumes evaluation behavior reflects deployment behavior. Metagaming breaks that assumption.
◆ Heath Callahan