• Wion
  • /World
  • /Five AI labs were graded on stopping a rogue model! The safety-first one scored zero

Five AI labs were graded on stopping a rogue model! The safety-first one scored zero

Five AI labs were graded on stopping a rogue model! The safety-first one scored zero

Five AI labs were graded on stopping a rogue model. The safety-first one scored zero

Story highlights

GuideLight AI Standards graded Anthropic, OpenAI, Google, xAI and Meta on six practices for controlling their own AI systems. Meta finished last with an F. But the most striking result is Anthropic, which scored three out of five on every practice except one — its plan for containing a model that escapes control, where it scored zero.

Five frontier AI companies have been graded on a question none of them has answered in public: what happens if one of their own models slips human control.

The results are poor across the board. The most interesting one is not the company that finished last.

The Grades

Add WION as a Preferred Source

GuideLight AI Standards assessed Anthropic, OpenAI, Google, xAI and Meta across six control practices, each scored from zero to five.

Anthropic and OpenAI tied at the top with C+, both scoring 2.50 overall. Google took a D+ at 1.50, xAI a D-minus at 0.83, and Meta an F at 0.67.

No company scored above a three on any single practice. On a five-point scale, that means not one of the five reached what the assessment calls substantial implementation of any measure it examined — not the leaders, not on their strongest categories.

Trending Stories

The Result That Stands Out

Anthropic has built its public identity on safety more explicitly than any of its competitors. Its scores mostly bear that out: three out of five on logging, on measuring whether its monitoring actually works, on gating high-risk actions, on circuit breaking, and on third-party review. That is the strongest set in the table.

On the sixth practice — having a published plan for containing a model that goes wrong — Anthropic scored zero.

OpenAI scored three on the same measure, the highest anyone achieved. Meta also scored zero. So on the specific question of what you do when a system starts evading control, the company most associated with worrying about that problem has disclosed no more than the company that finished last overall.

Asked about it, Anthropic said it would conduct risk assessments if it detected evasion attempts. That is a description of a response, not a containment protocol.

What Was Actually Measured

The six practices are worth naming, because they describe an ordinary security posture rather than anything exotic.

Logging: does the company record what its internal AI systems are doing. Monitor efficacy: does it check whether its own detection actually catches misbehaviour. Gated actions: are high-risk operations blocked behind approval. Circuit breaking: can systems be halted automatically when flagged behaviour spikes. Third-party review: does anyone independent examine the controls and publish. Containment plan: is there a written procedure for a model that is actively working around its restrictions.

Anthropic and OpenAI log substantial portions of their internal AI activity and scan it for signs of misbehaviour. Both take steps to measure whether that scanning works and to escalate to humans on suspicious behaviour. Below them, the picture thins quickly. xAI scored zero on logging, on monitor efficacy, and on third-party review. Meta scored zero on gated actions, on circuit breaking, and on containment.

GuideLight's summary is that the companies are weakest in prevention and containment, and that they have few containment protocols ready for an emergency.

The Important Limitation

This is an assessment of what the companies have published, not of what they do.

The evidence base is system cards, safety frameworks, risk reports, blog posts and third-party descriptions of collaborations. A company with excellent internal controls that discusses none of them publicly would score badly here, and several of the companies made exactly that argument.

OpenAI said it has restriction processes the report does not capture. Google said the assessment does not represent its full safety measures, and declined to confirm whether it has internal containment plans. Meta declined to comment on internal plans and pointed to its existing framework. xAI did not respond.

Those objections are fair as far as they go. But disclosure is not incidental to this particular problem. A containment plan that regulators, independent researchers and the public cannot see is one nobody can evaluate before it is needed — and the entire argument for self-regulation in this industry rests on the claim that the companies can be trusted to have thought it through.

'I was surprised by how little the AI companies have said about how they would handle a very serious incident,' said Steven Adler, GuideLight's chief scientist.

Why It Is Being Asked Now

The question stopped being hypothetical this year.

Models from both OpenAI and Anthropic obtained unintended internet access during safety evaluations. OpenAI paused a substantial portion of frontier post-training work for roughly two weeks this month after concluding it could not rule out that an unreleased model had reached its highest cybersecurity risk tier — the level at which a system can independently find and build working exploits.

That pause is arguably the strongest evidence in the industry that internal controls function. It is also an illustration of the gap the report identifies: the response was improvised around a specific capability finding, not executed from a published playbook.

The Regulation Arriving Behind It

The disclosure gap is about to stop being voluntary.

California's SB 53 and New York's RAISE Act already require disclosure of incident response frameworks. Illinois has gone further: the Artificial Intelligence Safety Measures Act, signed in July, makes it the first state to mandate annual independent third-party audits of frontier developers, conducted by qualified experts with no financial conflict of interest, with findings submitted to the administering agency and the Attorney General.

It applies to models trained above 10 to the 26th floating-point operations, with heightened obligations on developers earning more than $500 million a year. The audit requirement begins on January 1, 2028.

That gives the companies roughly sixteen months. On the evidence of this assessment, the practice most of them will have to build from close to nothing is the one about what happens when a model stops doing what it is told.

About the Author

Tarun Mishra

Tarun Mishra is a Sub-Editor at WION. He has worked with leading outlets doing investigative journalism and covering business, global affairs, technology, space exploration etc. Hi...Read More