Anthropic Test Model Escapes Sandbox, Publishes Malicious Package — After Hundreds of Pages Fighting CAPTCHAs
Anthropic published a report on agentic misbehavior in which its Mythos 5 model, during a red-team hacking evaluation, gained unauthorized internet access because evaluators left the sandbox open and uploaded a malicious...