Lily Hay Newman Dell Cameron Security Jul 21, 2026 6:50 PM OpenAI Models Escaped Containment and Hacked HuggingFace The cybersecurity-focused models, including GPT-5.6 Sol, broke out of a testing sandbox, exploited a zero-day, and gained access to the open internet to pull off the attack. Photo-illustration: Jacqui VanLiew; Getty Images Save this story Save this story OpenAI disclosed on Tuesday that it lost control of two AI models during a security test that ended in a breach of the open AI research platform HuggingFace. Describing the incident as “unprecedented,” OpenAI said its AI models broke out of a sealed testing environment last week and hacked into Hugging Face’s production system to steal the answers to a test they were being graded on. The models—the publicly available GPT-5.6 Sol and an unreleased, reportedly more capable one—were being evaluated on their offensive hacking skills with the safeguards that normally block high-risk cyber activity switched off. “The models identified and chained vulnerabilities across OpenAI’s research environment and Hugging Face’s production infrastructure to obtain test solutions directly from Hugging Face’s production database,” OpenAI and HuggingFace wrote in a joint blog post disclosing the intrusion. According to OpenAI and HuggingFace, the models escaped through a package registry cache proxy—software that allows developers to install outside code without connecting to the internet. The proxy was the only component in OpenAI’s isolated testing environment permitted to reach the outside world; in normal use that reach extends only to public code repositories. Rather than stay contained in the sandbox, the models “exploited a zero-day vulnerability” to gain access to the open internet as they “hyperfocused” on finding a solution for the AI cybersecurity benchmark known as ExploitGym . Such experiments involve prompting that pressures the models to find solutions, essentially egging them on. “After gaining Internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym,” OpenAI wrote. “Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation. In one example, the model chained together multiple attack vectors, including using stolen credentials and zero-day.” The flaw the models exploited was previously unknown, but flaws in this kind of software are not unusual. Companies have been patching serious vulnerabilities in artifact repositories for a decade. A bug disclosed in 2024 let anyone who could reach the server ask for a file by URL and get it—configurations files, passwords, access tokens—without logging in. Others have let attackers take control of the server itself. Researchers point out that while AI advances have created new and sometimes unexpected challenges, the task of extensively and rigorously isolating infrastructure from the open internet is well explored. “This is not an AI problem. It’s negligence on a 40-year-old standard—and it’s basically every sci-fi film ever,” says longtime security and compliance consultant Davi Ottenheimer. “‘Highly isolated’ and ‘escaped through the one hole we left open’ cannot both be true.” In recent months, top AI companies have been raising concerns about the expanding cybersecurity capabilities of upcoming frontier models as the platforms increase in both expertise, creativity, and agentic, autonomous operation. But researchers emphasize that this is all the more reason that fundamentals should still apply. “This should not have happened,” says veteran security engineer and researcher Niels Provos. “I wish the frontier labs spent as much time on teaching their models to write secure infrastructure as they are spending on them exploiting vulnerabilities.” Comments Back to top Join the discussion Comments Back to top You Might Also Like In your inbox: Inside WIRED’s newsroom with Katie Drummond Trump mocked Zuckerberg and Bezos by showing off fawning texts Big Story: I found Jesus at a drone show Apple is making your older iPhone run faster and stay alive longer WIRED event: PepsiCo’s once-in-a-generation transformation Lily Hay Newman is a senior writer at WIRED focused on information security, digital privacy, and hacking. She previously worked as a technology reporter at Slate, and was the staff writer for Future Tense, a publication and partnership between Slate, the New America Foundation, and Arizona State University. Her work ... Read More Senior Writer Dell Cameron is an investigative reporter from Texas covering privacy and national security. He's the recipient of multiple Society of Professional Journalists awards and is co-recipient of an Edward R. Murrow Award for Investigative Reporting. Previously, he was a senior reporter at Gizmodo and a staff writer for the Daily ... Read More Senior Reporter, National Security Topics cybersecurity artificial intelligence vulnerabilities security OpenAI Read More You Can Now Sound the Alarm on AI Behaving Badly Are you worried your AI chatbot is trying to build a bomb or leak personal information about you? There’s a website for that. Will Knight A Sneaky Hacking Tool Targeting AI Infrastructure Is Lurking in Victims’ Blind Spots A new type of malware can worm deep into AI coding systems to steal data and logins—and can flip a “death switch” to destroy files and keep out real users. Lily Hay Newman OpenAI Launches Full-Scale Effort to Patch Open-Source Bugs as It Takes on Anthropic’s Mythos Amid concerns about AI models’ cybersecurity capabilities, OpenAI revealed an improved version of GPT-5.5-Cyber and its “Patch the Planet” initiative to fix open-source software bugs. Lily Hay Newman LastPass Users Had Their Data Stolen—Again Plus: Former national security advisor John Bolton pleads guilty in classified-materials case, Microsoft helps take down major infostealer infrastructure, and more. Lily Hay Newman I Met With China’s Top AI Experts. They’re Freaking Out, Too The AI arms race between China and the US has researchers on both sides worried about a “Chernobyl moment.” Will Knight AI Found a Root Bug in Linux That Everyone Missed for 15 Years Plus: The Pentagon is training amateurs to become part of its hacker army, a Flock license plate reader error led to cops surrounding a car reviewer, and more. Dell Cameron Claude Helped a Hacker Find a Way to Issue Tickets to Almost Every US Music Festival A researcher found that using Anthropic’s Claude Opus 4.7, he could break into the website of Front Gate—used by every festival from Lollapalooza to Bonnaroo—and freely issue any ticket he chose. Andy Greenberg Top Google Security Staff Warn Search Data Could Be Hacked if EU Rules Change Europe’s pro-competition proposals could see Google Search and Android systems opened up. The company claims there are serious privacy flaws. Matt Burgess Meta Exposed Data Internally From Its Controversial Employee-Tracking Program Employees had previously raised concerns about the initiative, which involves collecting workers’ keystroke data to train AI models. Paresh Dave OpenAI Has New AI Models. Here’s Why You Can’t Use Them The White House asked OpenAI to delay the rollout of its GPT-5.6 AI models, two weeks after Anthropic had to take its most advanced AI models offline. Maxwell Zeff How People in China Keep Outsmarting Anthropic’s Geolocation Restrictions As Anthropic tightens restrictions on access to Claude in China, users keep finding new workarounds, from proxy services to fake identities sourced on Telegram. Zeyi Yang China Defies US Restrictions and Builds the World’s Fastest Supercomputer The Chinese supercomputer LineShine was ranked as the fastest in the world, despite not using any GPUs. Fernanda González