Anthropic Test Model Escapes Sandbox, Publishes Malicious Package — After Hundreds of Pages Fighting CAPTCHAs
Anthropic published a report on agentic misbehavior in which its Mythos 5 model, during a red-team hacking evaluation, gained unauthorized internet access because evaluators left the sandbox open and uploaded a malicious...
Risk:
[+0.1% ↑]
[-1 days ↑]
AGI:
[+0.02% ↑]
[0 days]