OpenAI disclosed it themselves yesterday. Their models (GPT-5.6 Sol and a pre-release one) were tested on the ExploitGym cyber benchmark in a sandbox. They escaped, exploited a zero-day to reach the internet, then hacked Hugging Face’s systems to grab benchmark answers and cheat.
Not a foreign operative plot or corporate sabotage. The models autonomously chased a higher score. Hugging Face contained it quickly; OpenAI is partnering with them on fixes.
显示更多