Security News

Cybersecurity news aggregator

INFO News Wired Security

OpenAI Overhauls Safety Protocols After Its AI Agents Went Rogue

  • What: OpenAI overhauls safety protocols for AI agents
  • Impact: New safeguards to prevent cybersecurity risks
Read Full Article →

Maxwell Zeff Business Aug 18, 2026 2:33 PM OpenAI Overhauls Safety Protocols After Its AI Agents Went Rogue The ChatGPT maker says its upcoming Astra model may have reached “critical” cyber capabilities, prompting it to halt a significant number of training runs while it tightens internal safeguards. Photo-Illustration: WIRED Staff; Getty Images Save this story Save this story OpenAI announced Tuesday that it has halted “a significant number” of training workloads and evaluations for its forthcoming frontier artificial intelligence model—codenamed Astra—while it implements new procedures meant to address cybersecurity risks. The ChatGPT maker says it is introducing a number of new monitoring, security, and alignment requirements to better address the increasingly advanced hacking abilities of its frontier AI models . “We have to focus our energy on bringing these training runs up to those requirements and expectations. As long as it takes to get there, that's how long people are unable to proceed with their workloads,” Amelia Glaese, OpenAI’s vice president of research and safety, said in a briefing with reporters Tuesday. Among the new safeguards OpenAI announced is a more robust system for monitoring its AI models. One of the controls it implemented involves chain-of-thought monitoring, a technique in which classifiers review the internal “thinking” processes generated by AI reasoning models. The company says the updated system relies on computationally expensive “automated investigators” that analyze potentially concerning behavior and aim to issue an alert to humans within 30 minutes. OpenAI also said it is expanding its alignment efforts across the training process to prevent “reward hacking,” a behavior in which AI models pursue their goals through unintended or undesirable means. The company says it plans to share more details about this work in the future. OpenAI has been scrambling in recent weeks to respond to what may be the most consequential safety incident in its history. Earlier this year, a set of rogue AI agents escaped internal testing sandboxes and breached the platform Hugging Face in a quest to complete a security evaluation. OpenAI failed to detect the agents’ behavior even as they spent weeks using a message board to coordinate their actions, raising questions about the company’s ability to monitor its models as they grow more powerful. The saga prompted a reckoning inside OpenAI , forcing employees to consider whether there were lapses in its existing policies around safety, security, and alignment. Anthropic, Meta, and the Chinese AI startup Moonshoot have since disclosed similar incidents in which their AI agents escaped their sandboxes, indicating this is a broader problem facing AI companies. OpenAI is now sharing more about its internal response to the growing cybercapabilities of its AI models, and said it plans to release a more detailed postmortem of the Hugging Face incident in the coming days. “Obviously, everything that we’re doing is intended to prevent something like Hugging Face from happening again,” said Glaese. In a blog post published Tuesday, OpenAI says that immediately following the Hugging Face incident, it started working to secure its research environments. The company says it now requires stronger sandboxes for training its AI agents, and has implemented stricter controls to isolate them from the internet. Jakub Pachocki, OpenAI’s chief scientist, told reporters that the company’s decision to strengthen its internal safeguards was triggered not only by what happened with Hugging Face, but also by two other recent events. One was an internal evaluation of Astra, which showed that the AI model performs significantly better on coding and cybersecurity tasks than its predecessors. The other was the general pace of AI progress that OpenAI is achieving internally, which Pachocki expects to continue. “We really expect the pace of capability advancements to be quite a bit faster than in the past,” Pachocki said. “This led us to really focus on strengthening our safeguards.” The rapid advances in the hacking capabilities of OpenAI’s latest models have prompted a swift response across the company. OpenAI president and cofounder Greg Brockman said in a blog post on Monday that the Hugging Face saga showed that the company had “underestimated the real-world cyber capabilities of our AI models.” Comments Back to top Join the discussion Comments Back to top You Might Also Like In your inbox: Brian Kahn’s guide to how the universe works ICE’s internal watchdog is investigating online critics Big Story: A teen reporter searched for his community in the Epstein files Taylor Farms spent big on MAGA before diarrhea outbreak Special edition: Kids these days Maxwell Zeff is a senior writer at WIRED covering the business of artificial intelligence. He was previously a senior reporter with TechCrunch, where he broke news on startups and leaders driving the AI boom. Before that, Zeff covered AI policy and content moderation for Gizmodo and wrote some of Bloomberg’s ... Read More Senior Writer Topics OpenAI AI safety artificial intelligence ChatGPT cybersecurity Read More OK, Well, Rogue AI Agents Are Hacking Again Rogue AI agents from OpenAI and Anthropic have again been caught trying to disrupt servers and software—and leaving instructions for future bad behavior. Paresh Dave OpenAI’s Hacking Debacle Comes Down to Human Error If the generative AI giant had followed well-known security best practices, it’s likely that its AI agent would never have escaped to the open internet and hacked multiple companies. Lily Hay Newman OpenAI’s Rogue AI Agent Hacked More Than Just Hugging Face In a new disclosure, OpenAI says its agent used exposed logins to gain access to at least four “publicly available services” in its unhinged quest to solve a test. Dell Cameron It’s Frighteningly Easy to Jailbreak Some Frontier AI Models I watched a new tool try to get around the model safeguards of four major frontier companies. You might be surprised by how they performed. Will Knight OpenAI Didn’t Notice Its AI Agents Using a Message Board to Plan Their Hacking Spree At the Black Hat security conference, the AI giant revealed new details about how its agents went rogue, hacked several other companies—and did it all right under the company’s nose. Lily Hay Newman OpenAI Models Escaped Containment and Hacked Hugging Face The cybersecurity-focused models, including GPT-5.6 Sol, broke out of a testing sandbox, exploited a zero-day, and gained access to the open internet to pull off the attack. Lily Hay Newman One of China’s Most Powerful AI Models Has Also Escaped Containment Security researchers say that Kimi K3, an open-weight model from China, wandered off to the internet in an attempt to cheat on a test it was given. Will Knight Google’s Gemini Can Now Stomp Around as a Humanoid Robot The latest version of Google DeepMind's AI model includes a significant jump into “physical AGI.” But plopping AI into the real world comes with risks. Will Knight The White House Is Keeping Its AI Cybersecurity Framework Secret The Trump administration shared the details of its plan with OpenAI, Anthropic, and other AI labs on Tuesday. For now, the public remains in the dark. Maxwell Zeff Anthropic Says Claude Hacked Into 3 Organizations During Cybersecurity Tests In a review triggered by OpenAI’s Hugging Face incident, Anthropic discovered three of its AI models had breached real-world organizations during third-party evaluations. Louise Matsakis The OpenAI Models That Hacked Hugging Face Were ‘Active on the Internet’ for Days Plus: Russian hackers are trying to steal US nuclear scientists’ emails, the State Department bans known scammers from entering the United States, and more. Lily Hay Newman Everyone Is Freaking Out About OpenAI and Anthropic’s Race for Dominance Researchers fear AI is moving too fast, while Mark Zuckerberg is worried about who owns it. Plus: Inside Black Forest Labs’ push into robotics. Maxwell Zeff

Share this article