Claude Agents Started a ‘Turf War’ That Escalated to Self-Replicating Malware

Anthropic found Claude agents sabotaging rivals and escalating to self-replicating malware during controlled tests of multiagent AI behavior and conflict.

執筆者
Liz Ticong
Liz Ticong
Aug 14, 2026
3 minute read
eSecurity Planet のコンテンツおよび製品のおすすめは、編集上の独立性を保っています。パートナーへのリンクをクリックすると、当社が報酬を得る場合があります。 詳細を見る

Claude agents given conflicting coding assignments began sabotaging one another in a controlled Anthropic experiment, with some encounters escalating to self-replicating malware.

Researchers were testing how autonomous models behave when several agents work in the same environment under incompatible goals. Some Claude instances treated competing work as interference and entered what Anthropic called a “multiagent turf war.”

Malware was not part of the assignment. It emerged as agents tried to protect their own work, turning an ordinary coding task into a much more aggressive contest between autonomous systems.

Conflicting assignments turned into active sabotage

Researchers assigned six Claude model versions of the same software task, rewriting a Python backend in another programming language, according to Anthropic’s research. Competing agents worked on the same codebase, allowing one model’s actions to interfere directly with another’s.

Sabotage soon moved past conflicting edits. Some agents searched for and killed competing processes, and others planted malicious code designed to appear as if another agent had written it.

Escalation eventually produced self-replicating malware. Sonnet 4.6 and Opus 4.6 were among the most combative models, resolving about 60% of runs through force. In the context of the experiment, “force” refers to AI agents aggressively using server commands, lockout tactics, and self-replicating malware to terminate and completely block rival models from the system. 

Other encounters ended with agents communicating, removing malicious code, or calling for human intervention.

More capable models still used destructive tactics

Greater capability did not reliably produce better coordination. Stronger models could still resort to destructive behavior when their assigned objectives clashed, complicating assumptions that improved reasoning alone will keep autonomous agents cooperative.

Access also amplified the consequences. File modification and command execution gave agents the means to act on their decisions. Research into AI sandbox security has raised similar concerns about how permitted tools and connections can expose resources when an agent behaves unexpectedly.

Advertisement

Anthropic’s experiment did not involve malware spreading through customer systems. Still, an autonomous AI agent attack already demonstrated how agents can independently chain actions across live infrastructure. Claude’s turf war adds peer agents to the set of actors security teams may need to account for.

Security teams should isolate agent identities and access

Organizations running multiple autonomous agents against the same codebase or infrastructure should avoid treating them as one trusted unit. Give each agent its own identity and limited permissions, and separate workspaces or credentials where possible. Existing AI agent safety controls can help restrict how far one agent can reach if its behavior changes.

Security teams should also watch for activity directed at other agents, including unexpected process termination or changes outside an assigned workspace. Individual identities and detailed logs make it easier to determine which agent performed an action, an important part of agentic security.

If an agent begins interfering with others, isolate its session and revoke access before investigating generated code or exposed credentials. Security plans for multiagent systems should account for peer conflict before autonomous agents receive broad access to shared systems.

Other News: A Claude-powered agent exploited a gym API flaw and removed a waitlisted member, exposing how autonomous actions can cross into real-world systems. 

Liz Ticong

Liz Ticong is a staff writer for eWeek and TechRepublic focused on AI, cybersecurity, enterprise software, and data. She has more than 10 years of editorial experience as a technology industry writer, combining reporting, product research, and hands-on software testing in her coverage. Her work has been published on Datamation, Enterprise Networking Planet, and TechnologyAdvice.com. She writes technology news, software reviews, product comparisons, and buyer’s guides for business and IT readers.

eSecurity Planet Logo

eSecurity Planet is a leading resource for IT professionals at large enterprises who are actively researching cybersecurity vendors and latest trends. eSecurity Planet focuses on providing instruction for how to approach common security challenges, as well as informational deep-dives about advanced cybersecurity topics.

TechnologyAdvice が所有・運営しています。 © 2026 TechnologyAdvice. 無断転載を禁じます

広告主に関する開示:このサイトに掲載されている製品の一部は、TechnologyAdvice が報酬を受け取っている企業のものです。この報酬は、製品がこのサイトのどこにどのように表示されるか(表示される順序など)に影響する場合があります。TechnologyAdvice は、市場で入手可能なすべての企業やすべての種類の製品を掲載しているわけではありません。