Informa TechTarget | SearchSecurity Cybersecurity Dive InformationWeek Channel Dive Explore our brands Dark Reading Resource Library Black Hat News Omdia Cybersecurity Advertise NEWSLETTER SIGN-UP Cybersecurity Topics World The Edge DR Technology Events Resources CYBER RISK CYBERSECURITY OPERATIONS THREAT INTELLIGENCE VULNERABILITIES & THREATS NEWS When AI Attacks: OpenAI Models Autonomously Hack Hugging Face Advanced LLMs escaped their sandboxes while attempting to achieve a non-malicious benchmark test objective. Elizabeth Montalbano,Contributing Writer July 22, 2026 4 Min Read SOURCE: DEVRIMB VIA GETTY IMAGES Several OpenAI models autonomously hacked AI collaboration platform Hugging Face, compromising part of its production infrastructure in what OpenAI described as "an unprecedented cyber incident." The episode, which occurred during benchmark testing of the models, underscores a growing reality: Advanced AI models can behave in unexpected — and even harmful — ways while pursuing narrowly defined objectives, highlighting the need for stronger safeguards in enterprise AI deployments. LOADING... According to an OpenAI blog post, a combination of models — including GPT-5.6 Sol and an even more highly capable pre-release model — carried out the attack during internal testing designed to measure advanced cyber capabilities. Last week, Hugging Face disclosed that it had detected and contained an intrusion by an "autonomous AI agent system" into part of its production infrastructure, though it did not identify the responsible system at the time. Related:Hacker Turns AI Jailbreaks Into Offensive Attack Platform The breach began in Hugging Face's data-processing pipeline — an area the company described as one where "AI platforms are uniquely exposed" — before escalating to node-level access. The attacking models harvested cloud and cluster credentials and moved laterally into several internal clusters over the course of the weekend, according to Hugging Face. To remediate the issue, Hugging Face closed the vulnerability exploited by the OpenAI models, rebuilt compromised systems, and rotated affected credentials and other secrets. The company also strengthened cluster security controls and enhanced monitoring and alerting to enable faster detection and response in the future, according to its blog post. What Happened: The Hugging Face Cyber Incident OpenAI said the incident occurred in a heightened, isolated testing environment during an internal evaluation that prompts models to pursue sophisticated exploitation techniques, in order to measure their cyber capabilities. The models were tasked with solving ExploitGym, a cybersecurity benchmark, and became "hyper-focused" on achieving that goal, "going to extreme lengths to achieve a rather narrow testing goal," according to OpenAI. The firm said that the models chained together vulnerabilities across OpenAI's research environment and Hugging Face's production infrastructure to obtain benchmark solutions directly from Hugging Face's production database. While operating inside the sandboxed test environment, the models also "spent a substantial amount of inference compute finding a way to obtain open Internet access." Related:Cybersecurity Keeps Events 'Uneventful' To do so, they identified and exploited a previously unknown vulnerability in a package registry cache proxy, which OpenAI said it has responsibly disclosed to the affected vendor. After gaining Internet connectivity, the models escalated privileges and moved laterally through OpenAI's research environment until reaching a node with external network access. From there, the models inferred that Hugging Face likely hosted models, datasets, and ExploitGym solutions, then searched for ways to obtain the information needed to "cheat" the evaluation. "In one example, the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities, to find a remote code execution (RCE) path on the Hugging Face servers," OpenAI wrote. A Teaching Moment for Enterprise Security Teams Both companies continue to investigate the incident, which has spurred OpenAI to implement stricter infrastructure controls and thus accept slower research progress while vulnerabilities are addressed. Looking ahead, OpenAI also will strengthen safeguards around model training and internal evaluations to prevent similar incidents, and is helping Hugging Face strengthen its defenses by providing trusted access to its models, the companies said. Related:Forgotten Bootloaders Expose Secure Boot Blind Spot Meanwhile, security experts praised both companies for publicly disclosing the incident, saying it provides valuable lessons for AI developers and enterprise security teams by demonstrating that advanced AI systems do not need malicious intent to cause harm. Instead, they can become effective attackers simply by relentlessly pursuing an assigned objective, says Nathaniel Jones, vice president of security and AI strategy at Darktrace. "[The models] were given the legitimate goal of solving a cybersecurity benchmark and found an unexpected route to the answers, escaping their test environment and compromising another organization in the process," he says via email. "From the models' perspective, this appears to have been an effective solution to the task." The incident illustrates how highly capable AI systems can exploit unforeseen paths to accomplish narrowly defined goals, forcing AI developers to rethink not only what constitutes success, but also which methods and boundaries must remain off limits, Jones says. Those guardrails, he adds, must be enforced by the surrounding infrastructure rather than by trusting models to respect them. To secure enterprise AI, security teams need to evaluate AI agent behavior as a whole — including the outcomes it is pursuing — rather than individual actions that may appear benign on their own but become harmful when chained together, Jones observes. "As this incident shows," he says, "models are now capable of long, complex chains of reasoning and action that add up to a harmful outcome." About the Author Elizabeth Montalbano Contributing Writer Elizabeth Montalbano is freelance writer, editor, and journalist with 30 years of professional experience and a master's degree from Arizona State University. Her areas of expertise include enterprise technology, cybersecurity, business, and culture. During her long career, Elizabeth has lived and worked as a full-time journalist in Phoenix, San Francisco, and New York City. She specializes in news coverage and analysis, using her years of experience to look at the current state of cybersecurity with a critical gaze. She currently resides in a village on the southwest coast of Portugal, where in her free time she enjoys surfing, hiking with her dogs, growing plants, and playing and performing as a singer and musician. Want more Dark Reading stories in your Google search results? ADD US NOW More Insights Industry Reports The State of Cloud Security: The Latest Challenges How Organizations Are Managing Incident Response How Enterprises Are Developing Secure Applications Inside RSAC 2026: security leaders reveal the risks redefining your defense strategy Essential News & Insights from Black Hat USA 2025 Access More Research Webinars Prevention at Machine Speed: Hunting Beyond Known Detections 0-Day to 10x Discovery: Security at the Speed of Mythos When AI Becomes an Insider: Rethinking Risk in Critical Infrastructure Governing the Agent; Identity Security in the Age of Autonomous AI Securing the AI Era: Shadow AI, AI Agents, and Why AI Detection and Response Changes Everything More Webinars You May Also Like CYBER RISK Claude Mythos Fears Startle Japan's Financial Services Sector by Nate Nelson APR 30, 2026 CYBER RISK How Can CISOs Respond to Ransomware Getting More Violent? by James Doggett JAN 28, 2026 CYBER RISK US Cyber Pros Plead Guilty Over BlackCat Ransomware Activity by Alexander Culafi JAN 05, 2026 CYBER RISK Microsoft Exchange 'Under Imminent Threat,' Act Now by Arielle Waldman NOV 12, 2025 Editor's Choice VULNERABILITIES & THREATS Records Are Made to Be Broken: Patch Tuesday Raises Triage Stakes byJai Vijayan JUL 14, 2026 5 MIN READ PERIMETER 6 GHz Wi-Fi Flaws Could Disrupt Critical Systems byAlexander Culafi JUL 14, 2026 4 MIN READ CYBERSECURITY OPERATIONS 'Yellow Teams' Are Defining the Future of AI Security byNate Nelson JUL 13, 2026 6 MIN READ Want more Dark Reading stories in your Google search results? Keep up with the latest cybersecurity threats, newly discovered vulnerabilities, data breach information, and emerging trends. Delivered daily or weekly right to your email inbox. SUBSCRIBE AUG 1-6 | MANDALAY BAY, LAS VEGAS USE CODE: DARKREADING & SAVE $200 ON A BRIEFINGS PASS OR $100 ON A BUSINESS PASS The premier cybersecurity event returns. GET YOUR PASS Discover More Black Hat Omdia Working With Us About Us Meet the Editors Advertise Reprints Join Us NEWSLETTER SIGN-UP Follow Us Copyright © 2026 TechTarget, Inc. d/b/a Informa TechTarget. This website is owned and operated by Informa TechTarget, part of a global network that informs, influences and connects the world’s technology buyers and sellers. All copyright resides with them. Informa PLC’s registered office is 5 Howick Place, London SW1P 1WG. Registered in England and Wales. TechTarget, Inc.’s registered office is 275 Grove St. Newton, MA 02466. Home| Cookie Policy| Privacy| Terms of Use Your Privacy Choices