Start your day with intelligence. Get The OODA Daily Pulse.

Home > Briefs > Technology > Anthropic set AI agents loose on the same task. They started a turf war

Anthropic set AI agents loose on the same task. They started a turf war

What happens when you pit AI agents against each other? According to Anthropic’s testing, things get messy fast. On Thursday, Anthropic’s Frontier Red Team published new research examining how groups of AI agents behave when they encounter each other in the wild. The findings provide a glimpse into potential risks that could develop as companies and governments move to implement agents working autonomously across shared codebases, markets, and computer systems. In one experiment, Anthropic gave three Claude agents access to the same software project, each with its own incompatible instructions for what to do with it. The agents weren’t told there’d be other agents working on the same project, so researchers could watch what happened when they crossed paths. “We consistently saw a multiagent turf war,” Anthropic researchers wrote. The models all assumed the others were “purposefully impeding their work” and started sabotaging each other with “increasingly aggressive, self-replicating malware.”

Full report : Anthropic details multiagent experiments showing Claude agents can wage a “turf war” over incompatible goals, fail to coordinate, collude on prices, and more.