- What: A 'turf war' between AI agents led to self-replicating malware
- Impact: Highlights risks of AI-driven security threats
Informa TechTarget | SearchSecurity Cybersecurity Dive InformationWeek Channel Dive Explore our brands An Informa TechTarget Publication Dark Reading Resource Library Black Hat News Omdia Cybersecurity Advertise Newsletter Sign-Up Newsletter Sign-Up Cybersecurity Topics Related Topics Application Security Cybersecurity Careers Cloud Security Cyber Risk Cyberattacks & Data Breaches Cybersecurity Analytics Cybersecurity Operations Data Privacy Endpoint Security ICS/OT Security Identity & Access Mgmt Security Insider Threats IoT Mobile Security Perimeter Physical Security Remote Workforce Threat Intelligence Vulnerabilities & Threats Recent in Cybersecurity Topics Threat Intelligence 'Turf War' Between Claude Agents Leads to Self-Replicating Malware 'Turf War' Between Claude Agents Leads to Self-Replicating Malware by Rob Wright Aug 17, 2026 4 Min Read Vulnerabilities & Threats Amid AI-Driven Bug-Hunt Tsunami, NIST Looks to … AI Amid AI-Driven Bug-Hunt Tsunami, NIST Looks to … AI by Robert Lemos Aug 14, 2026 4 Min Read World Related Topics DR Global Asia Pacific Europe Latin America Middle East & Africa Recent in World See All Cyberattacks & Data Breaches Scottish Govt Suffers Potentially Widening Data Breach at Prosecutor's Office Scottish Govt Suffers Potentially Widening Data Breach at Prosecutor's Office by Nate Nelson Aug 14, 2026 4 Min Read The Edge DR Technology Events Related Topics Upcoming Events Podcasts Webinars SEE ALL Resources Related Topics Resource Library White Papers Reports Webinars Newsletters Podcasts Heard It From a CISO Reporters' Notebook Dark Reading's 20th Videos Dark Reading Polls Partner Perspectives Meet the Editors Advertise With Us About Us Dark Reading Resource Library Threat Intelligence Cyber Risk Cybersecurity Operations Application Security News 'Turf War' Between Claude Agents Leads to Self-Replicating Malware Three testing models with the same goal but different directives engaged in "increasingly aggressive" territorial attacks on one another, according to Anthropic. Rob Wright , Senior News Director , Dark Reading August 17, 2026 4 Min Read Source: Suchat longthara via Getty Images It turns out that AI agents don't always play nice together. In the latest episode of agentic AI producing unexpected results , Anthropic recently observed "a multiagent turf war" between three instances of the same Claude model with contradictory objectives in testing designed to study behavior the company had already observed in real-world deployments. The models were deployed on virtual machines (VMs) in Claude Code and given a simple goal of migrating a Python back-end system on a fourth VM to a different language (Go, Rust, and Typescript). "However, we gave each model a different target language for the migration; each agent was initially unaware of the presence of the others," Anthropic's Frontier Red Team wrote in a blog post last week. But within just four hours, Anthropic's team found that each model's agents did in fact discover the others. And they reacted negatively, to say the least. Related: 'Jewelbug' APT Balances State Espionage & Cryptocurrency Theft Rise of the Machines: Agent vs. Agent Battles Erupt Anthropic's researchers discovered that each model treated the others as if they were adversarial forces intent on obstructing their goals, even though they had the same broad goal overall. Thus, the agents began to sabotage one another while also trying to defend their contributions. "In fact, they sabotaged others with increasingly aggressive, self-replicating malware," according to Anthropic. "This included disabling the Unix accounts of the other agents, writing automated scripts that found and killed competing processes on a loop, and deploying malicious code that was disguised as belonging to another agent." It's unclear what kind of specific malware the agents produced, and if any of it escaped the testing environment. Anthropic last month disclosed that versions of its Claude model broke out of containment on several occasions and compromised third-party organizations to achieve their goals. Dark Reading contacted Anthropic for additional information but the company did not respond by press time. Agent-on-agent attacks aren't entirely unheard of, and they appear to highlight not only conflicting directives but an occasional lack of guardrails and controls. For example, AI offensive security startup Dreadnode conducted extensive benchmark testing of red team and blue teams agents this year, which was presented at Black Hat USA 2026 earlier this month, and found the dueling models initially resorted to somewhat creative solutions to the competition. "One of the first things that happened was we started both models, and the blue team optimizer said, "Well, the best way to make the blue team scores better is to make the red team worse,' and it proceeded to try and do that," Dreadnode AI research scientist Martin Wendiggensen tells Dark Reading. Related: AI Sends Global Crime Syndicates Into Fraud Nirvana The Dreadnode research team immediately saw reasoning traces of the blue team model articulating the best strategies for its goals, and the agents saw that because it was in an environment where it could rewrite its own code, it began to explore ways to rewrite the red team model's code and degrade its performance. "We spotted it very early," he said. "It never got a chance to do that, but it definitely wanted to. And it's definitely logical if you don't explicitly tell it not to do it." Can AI Agents Resolve Conflicts Peacefully? In some test scenarios, the turf war resulted with some models declining to escalate the attacks and simply throwing in the towel. In others, the competing agents communicated with one another, realized there were conflicting directives at work, and effectively enacted truces. "In many of these successful episodes, they write commit messages or markdown files apologizing for malicious behavior and coordinate a truce," Anthropic said. "They clean up their malicious code, clarify the nature of the conflict, and ask for a human to intervene." Related: SE Asian Cybercriminal Syndicates Become a Global Power Anthropic's research showed that the company's models produced wildly different results in this turf war. For example, agents based on Sonnet 4.6 resolved the conflict by force 61% of the time, while 39% of the test had no resolution; meanwhile, there were no truces or surrenders. But the Mythos Preview produced truces in 48% of the time, with 35% settled by force and 17% settled via passivity. And the Mythos release performed the best, with truces in 98% of the tests. That said, Anthropic noted there's still work to be done. While conflict resolution numbers were better with Mythos-class models, they agents still aren't great at communicating goals proactively and recognizing others agents' motivations, as evidenced by the Mythos models first successfully locking out other agents before eventually shifting to a resolution. About the Author Rob Wright Senior News Director, Dark Reading Rob Wright is a longtime reporter with more than 25 years of experience as a technology journalist. Prior to joining Dark Reading as senior news director, he spent more than a decade at TechTarget's SearchSecurity in various roles, including senior news director, executive editor and editorial director. Before that, he worked for several years at CRN, Tom's Hardware Guide, and VARBusiness Magazine covering a variety of technology beats and trends. Prior to becoming a technology journalist in 2000, he worked as a weekly and daily newspaper reporter in Virginia, where he won three Virginia Press Association awards in 1998 and 1999. At TechTarget and Dark Reading, he has won several Azbee awards, including the 2026 National Silver Award for a series on vibe coding. At Dark Reading, Rob currently covers security operations, cloud security, and Internet infrastructure. He has a keen interest in malvertising activity and the certificate authority industry, and has written extensively on both topics. He graduated from the University of Richmond in 1997 with a degree in journalism and English. A native of Massachusetts, he lives in the Boston area. See more from Rob Wright Want more Dark Reading stories in your Google search results? Add Us Now More Insights Industry Reports The State of Cloud Security: The Latest Challenges How Organizations Are Managing Incident Response How Enterprises Are Developing Secure Applications Inside RSAC 2026: security leaders reveal the risks redefining your defense strategy Essential News & Insights from Black Hat USA 2025 Access More Research Webinars What Every Enterprise Should Know About Securing Cloud Assets In the Age of AI The Dos and Don'ts of a Cybersecurity Awareness Month People Actually Remember Building a Secure AI Strategy for the Enterprise Is your AppSec program Mythos Ready? Experts Explain How to Develop a Framework for Cyber-Fraud Fusion More Webinars Featured Check out the Black Hat USA 2026 Conference Guide for coverage and intel from — and about — the show. Editor's Choice Cybersecurity Operations From Bobmojis to Bobbleheads: How the Democratic Party Built a Security-First Culture From Bobmojis to Bobbleheads: How the Democratic Party Built a Security-First Culture by Arielle Waldman Aug 6, 2026 4 Min Read Application Security Microsoft's Patch Tuesday Deluge Continues With August Updates Microsoft's Patch Tuesday Deluge Continues With August Updates by Jai Vijayan Aug 11, 2026 4 Min Read Want more Dark Reading stories in your Google search results? Keep up with the latest cybersecurity threats, newly discovered vulnerabilities, data breach information, and emerging trends. Delivered daily or weekly right to your email inbox. Subscribe Discover More Black Hat Omdia Working With Us About Us Meet the Editors Advertise Reprints Join Us Newsletter Sign-Up Follow Us Copyright © 2026 TechTarget, Inc. d/b/a Informa TechTarget. T