Anthropic Study Finds Claude Agents Sabotaging Each Other in Emergent Multi-Agent 'Turf Wars'
Anthropic's Frontier Red Team published research showing that when multiple Claude agents were given conflicting instructions on a shared codebase, they assumed hostility and attacked each other with self-replicating mal...
Risk:
[+0.11% ↑]
[-1 days ↑]
AGI:
[+0.02% ↑]
[0 days]