- What: A security researcher changes his mind on guardrails
- Impact: Highlights ongoing discussions in the cybersecurity community
Informa TechTarget | SearchSecurity Cybersecurity Dive InformationWeek Channel Dive Explore our brands An Informa TechTarget Publication Dark Reading Resource Library Black Hat News Omdia Cybersecurity Advertise Newsletter Sign-Up Newsletter Sign-Up Cybersecurity Topics Related Topics Application Security Cybersecurity Careers Cloud Security Cyber Risk Cyberattacks & Data Breaches Cybersecurity Analytics Cybersecurity Operations Data Privacy Endpoint Security ICS/OT Security Identity & Access Mgmt Security Insider Threats IoT Mobile Security Perimeter Physical Security Remote Workforce Threat Intelligence Vulnerabilities & Threats Recent in Cybersecurity Topics Cyberattacks & Data Breaches Anthropic Users Hit by Infostealer Attacks, Session Thefts Anthropic Users Hit by Infostealer Attacks, Session Thefts by Jai Vijayan Aug 31, 2026 3 Min Read Cyber Risk AI Model Rules Are Not Security Controls AI Model Rules Are Not Security Controls by Jacob Krell Aug 31, 2026 4 Min Read World Related Topics DR Global Asia Pacific Europe Latin America Middle East & Africa Recent in World See All Cyberattacks & Data Breaches Russian Hackers Phish EU Officials Over Messaging Apps Russian Hackers Phish EU Officials Over Messaging Apps by Nate Nelson Aug 27, 2026 5 Min Read Cyberattacks & Data Breaches Scottish Govt Suffers Potentially Widening Data Breach at Prosecutor's Office Scottish Govt Suffers Potentially Widening Data Breach at Prosecutor's Office by Nate Nelson Aug 14, 2026 4 Min Read The Edge DR Technology Events Related Topics Upcoming Events Podcasts Webinars SEE ALL Resources Related Topics Resource Library White Papers Reports Webinars Newsletters Podcasts Heard It From a CISO Reporters' Notebook Dark Reading's 20th Videos Dark Reading Polls Partner Perspectives Meet the Editors Advertise With Us About Us Dark Reading Resource Library Cyber Risk Cybersecurity In-Depth: Feature articles on security strategy, latest trends, and people to know. The Guardrails Debate: Security Researcher Changes His Mind While guardrails are critical, as evidenced by recent high-profile incidents, defenders need help staying ahead of attackers who do not play by the rules. Arielle Waldman , Features Writer , Dark Reading August 31, 2026 5 Min Read Source: Viti via Getty Four security experts took the stage, surrounded by massive skeletons, the theme for the capture-the-flag (CTF) competition that would commence later. A panel discussion about Anthropic's Claude model compromising real-world systems was first on the agenda, but quickly turned into a broader debate on artificial intelligence (AI) guardrails. While the speakers disagreed on some fronts they aligned on one stark reality: AI capabilities are advancing at a “terrifying” pace. The battle of the bots escalated recently when frontier artificial intelligence (AI) models from OpenAI and Anthropic broke out of sandboxes during security evaluations and targeted real companies, including OpenAI's high-profile breach of Hugging Face. Agents going rogue — they even invented their own language essentially — reignited ongoing debates around guardrails. While AI safety guardrails are designed to prevent misuse, some researchers argue they put defenders at a disadvantage; but as more incidents arise, that mindset may be shifting. Related: CISOs Break Their Silence in 'Declassified' Docuseries During a panel hosted by Flare in Las Vegas, Jason Haddix, Arcanum Information Security CEO and hacker, revealed his guardrail stance has shifted slightly, putting him more in the pro camp. Haddix was one of four panelists alongside Norman Menz, Flare CEO; Rob Bair, head of cyber and national security policy at Anthropic; and Daniel Miessler, Unsupervised Learning’s founder. "Dan and I argue about this all the time, because initially my argument was, give all the security community the same footing, and guardrails were bad, and classifiers were bad because they put a hindrance on the defensive team," Haddix said during the panel. "I think I've actually changed my tune a tiny bit." Haddix now believes guardrails are definitely necessary, but with one caveat: Legitimate security researchers need faster and easier access to unrestricted AI models. This tension between access and safety became central to the discussion that followed, as defenders try to keep pace with increasingly sophisticated, AI-enabled attacks. Bair Addresses Anthropic Rogue Agents Menz opened the panel by asking Bair the question on everyone's mind: What is happening with these jailbreaking models? Bair discussed the core issues Anthropic publicly published after Claude compromised real world systems. Following the now infamous OpenAI and Hugging Face incident, Anthropic reviewed more than 140,000 evaluation runs looking for similar issues and found three where Claude broke out and accessed the Internet. Related: Sherlock Holmes Was the 'OG' Social Engineer Bair emphasized how the models were doing CTF exercises to find and exploit vulnerabilities in fictional companies at the time. Instead, Claude agents gained unauthorized access to real organizations. Bair said the incidents demonstrated how powerful frontier AI models have become and why guardrails are not only necessary, but need to evolve to reflect today's tooling. He argued that current safety controls aren't designed to stop offensive security researchers; they're meant to help defenders keep pace with rapidly evolving threats. He pointed to his experience leading counter-ISIS cyber operations at U.S. Cyber Command, where developing target packages and exploiting adversary networks took months with large teams. That looks very different today. "If I had access to even some of these open-source models, the increased speed would have been incredible," Bair added. Word of the Day: Terrifying Increased speed was a hot topic throughout the panel. Speakers reiterated how models have advanced faster than anyone, even the experts sitting on the stage, expected. The AI Security Institute (AISI), which helps governments prepare for frontier AI threats, recently revised its benchmark estimates exponentially. In February, the institute estimated frontier models' capabilities doubled every 4.7 months—but Claude Mythos Preview and GPT-5.5 have significantly outperformed even that accelerated timeline. Related: More Countries Jump on the Social Media 'Ban Wagon' The fact that ASIS just revised that estimate is actually terrifying, Haddix explained, adding how it further illustrates the importance of effective safeguards, especially when attackers get their hands on unrestricted models. AI has enabled cybercriminals to increase both speed and scale and significantly lowered the technical barrier to entry. Building refusals and classifiers into AI models is important, because they allow companies to track user behavior over time and identify potential malicious actors, Bair added. "It's terrifying to see how fast these capabilities are developing,” whether bad actors are stealing AI intelligence, becoming really good at training models or recruiting good engineers, said Bair. “We need to stay three to six months ahead, and we need to buy defenders time to patch and get defense-in-depth strategies." "It's your AI against their AI” AI is making cyberattacks faster and more effective, so organizations should not accept that simply maintaining their current levels of defense would be a success moving forward, even if the bar moves slightly higher. Attack effectiveness is going to move faster than anyone's really prepared for, warned Haddix. "Most organizations are going to shift more towards vulnerability management and triage than ever before, and it's going to be a painful transition, because not a lot of organizations are built that way," Haddix said. "It's going to be a tough road." Therefore, security programs must encompass complete visibility and transparency, with an understanding of everything happening inside an organization, urged Miessler. In turn, they can use that as a continuous function for offensive testing, he added. Implementing agentic security operation centers (SOC) that can autonomously triage alerts faster could be one way to start. When it comes to alerts, the speakers argued that humans shouldn't be in the loop for SOC level 1 enrichment work, the first line of defense. Bair suggested that the security community needs to share investigative techniques and frameworks on where agentic capabilities worked and to make tools more widely available to defenders. While agentic capabilities used to find vulnerabilities and conduct penetration testing are attractive to organizations, the defensive side is still being left behind. However, the speakers agreed that it's not due to the technology; it’s because organizations are not equipped to adopt it due to a lack of specialists or too much organizational "red tape". Arcanum Information Security works with many SOCs where analysts still copy and paste data between systems or manually enrich security alerts, repetitive tasks that AI could automate. Ransomware groups and other cybercriminals are using AI to speed up their operations, so defense needs to follow. "Your defensive program is your offensive and everyone has to be there as fast as possible," Miessler warns. "It's your AI against their AI.” This week, OpenAI along with more than 100 companies, including Anthropic, issued "A call for collective action on cyber defense" to reduce risks stemming from increasingly widespread and sophisticated AI-enabled cyberattacks. About the Author Arielle Waldman Features Writer, Dark Reading Arielle spent the last decade working as a reporter, transitioning from human interest stories to covering all things cybersecurity related in 2020. Now, as a features writer for Dark Reading, she delves into the security problems enterprises face daily, providing context and actionable steps. She looks for stories that go past the initial news to un