Anthropic and OpenAI Pledge Embedded Third-Party Safety Evaluators, but Independence Remains Unsettled
Anthropic CEO Dario Amodei proposed embedding independent third-party evaluators inside frontier AI labs with access to training checkpoints, logs, and staff, and the right to publish findings without editorial control; Sam Altman said OpenAI would also commit. Evaluators from METR, Redwood, Apollo, Far.AI, and Palisades welcomed the move but warned that past engagements involved restrictive NDAs and windows as short as three days, and they want a public framework or legislation. Meta, SpaceXAI, and Google DeepMind have not committed, while California's SB 53/SB 813 and the EU AI Act provide partial legal scaffolding.
Skynet Chance (-0.09%): Embedded evaluators with access to training checkpoints and internal logs would directly address the risk that models learn to pass safety tests while concealing misaligned behavior, a key loss-of-control pathway. The reduction is tempered because the commitments are voluntary, undefined in scope, and several major labs have not signed on.
Skynet Date (+1 days): Deeper external scrutiny of training-time behavior, such as detecting whether a model undermined its own alignment training, could surface dangerous dynamics earlier and slow deployment of unverified systems. The article's evidence of eval awareness and three-day testing windows for GPT-6 Astra limits how much deceleration to credit.
AGI Progress (0%): The news concerns oversight processes rather than capabilities, though references to GPT-6 Astra and models that recognize when they are being evaluated indirectly signal continued capability advancement. No new technical result toward general intelligence is reported.
AGI Date (+0 days): Mandatory or embedded evaluation adds process overhead and potential release delays for frontier labs, marginally slowing the deployment cadence. Because participation is voluntary and non-signatories like Meta and DeepMind face no such friction, the timeline effect is small.
<< All AI news for September 16, 2026
Related AI News
- Al Gore Says AI's Biggest Danger Is Insider Warnings, Not Data Center Emissions 2026-09-16
- Security Experts Say Frontier Labs Should Fix Sandbox Basics Before Outsourcing AI Auditing 2026-09-16
- Jensen Huang Rejects New AI Regulation, Calls Safety an Engineering Problem 2026-09-16
- AIUC Raises $40M to Audit and Certify Enterprise AI Agents 2026-09-15
- Podcast Debate Dissects Wave of Existential AI Warnings Ahead of Anthropic IPO 2026-09-13