Frontier AI Labs Have Few Public Plans to Contain a Rogue Model, Study Finds
An independent Guidelight assessment graded five frontier AI labs on six control practices and found few publicly documented plans for containing a rogue model. The scores measure public disclosure, not private preparedness, but only OpenAI cleared even a partial mark on containment.
Guidelight AI Standards, a new safety-standards group, published its first assessment of how frontier AI labs would keep control of their own models, and the scorecard is blunt: across five leading companies and six basic control practices, no lab scored higher than a 3 out of 5 on anything. On the specific question of whether a company has a documented plan to contain a rogue model, only OpenAI cleared what Guidelight calls substantial partial implementation. Meta and Anthropic scored zero. The whole exercise rests on one caveat that defenders should read carefully: the grades measure what each lab has disclosed in public, not what it may be doing in private.
What Guidelight Graded, and How
The report, "AI Control: An Assessment of Frontier Practices," comes from Guidelight AI Standards and its chief scientist Steven Adler, a former OpenAI safety researcher. It scored Anthropic, Google, Meta, OpenAI, and xAI against six foundational practices from Guidelight's Control standard: logging what internal AI systems do, measuring how well that monitoring works, gating high-risk actions behind a monitor, circuit-breaking (halting systems after a surge of flagged misbehavior), independent third-party review, and a containment plan for a misaligned model.
The methodology is the part to underline. Guidelight worked only from public materials (system cards, safety frameworks, risk reports, blog posts, and third parties' write-ups) and graded each lab on what it could verify from the outside. As the report states plainly, its assessment "is based only on publicly available information," so a low mark reflects a lack of publicly documented plans, not proof that a safeguard is missing. Guidelight defines a containment plan precisely: a "pre-specified plan, triggered when the AI is detected trying to subvert control, which covers what permissions to revoke from the model, who the model may continue operating for, under what constraints, and when to take it fully offline."
"I was surprised by how little the AI companies have said about how they would handle a very serious incident if their model did escape their control in some sense," Adler told TechCrunch, which first reported the study.
Detection Is the Strong Suit. Containment Is the Gap.
No company came out looking finished. On Guidelight's 0-to-5 scale, no lab exceeded a 3 (substantial partial implementation) on any single practice, and most scores landed at 2 or below. The overall grades: Anthropic and OpenAI tied at C+ (2.50), Google at D+ (1.50), xAI at D- (0.83), and Meta at F (0.67).
The pattern underneath the grades matters more than the letters. Labs are relatively strong on detection and outside review. Anthropic and OpenAI both log large portions of internal AI activity and scan it for misbehavior, and four of the five (all but xAI) took part in METR's first Frontier Risk Report. They are weakest on prevention and containment. On the containment-plan practice specifically, OpenAI scored a 3, Google a 2, xAI a 1, and Meta and Anthropic both scored 0. Guidelight's summary line is the one to sit with: the best public evidence is that companies have "few containment protocols ready for an emergency."
That Anthropic lands at zero on containment despite its public emphasis on safety is the report's most counterintuitive result, and it is worth being clear about why. Guidelight found that Anthropic's August risk report does not mention limiting a model's deployment as a possible outcome of its process for investigating control incidents. That is a documentation gap, not a demonstrated failure to act.
Why "No Public Plan" Is Not the Same as "No Plan"
This is where an evenhanded read is essential, and where the labs pushed back. Several told TechCrunch that the assessment misses their internal work. An OpenAI spokesperson said the company has "a process for requiring restricting permissions, pausing workloads, limiting deployment, or taking the model fully offline, and [has] applied it." A Google spokesperson said the report does not capture the full scope of its safety measures. Anthropic said that if it detected a model trying to subvert human control, it would run a risk assessment to decide whether containment was the appropriate response. Meta pointed to an existing risk framework, and xAI did not respond in time.
There is also a plausible reason to stay vague that has nothing to do with negligence. Lily Li, a privacy and AI lawyer who founded Metaverse Law, told TechCrunch that firms may hold back specifics for legal, not just competitive, reasons: an overly precise public promise you then fail to keep "could form the basis of an unfair and deceptive marketing claim." Some of the silence Guidelight is measuring may be lawyers rather than missing plans. The finding is about disclosure and transparency. It is not proof that any given lab is unprepared.
The Regulators Are Already Forcing the Question
Transparency is exactly what a wave of new law is starting to demand. California's SB 53, in effect this year, requires large frontier developers to publish frameworks explaining how they respond to critical safety incidents and manage the risk of models circumventing oversight. New York's RAISE Act, with similar requirements, takes effect in January. And last month a bipartisan federal bill, the AI Kill Switch Act, proposed requiring major developers to build and maintain technical mechanisms to shut down a rogue model. "A kill switch is the bare minimum for today's models," Connor Leahy, U.S. executive director of the nonprofit ControlAI, told TechCrunch.
This is the same throughline running through OpenAI's own recent moves, which we have tracked closely: its overhaul of safety protocols after two test models went rogue and the subsequent two-week pause on frontier reinforcement-learning training, which put a 20% price tag on watching its own systems. Read together, Guidelight's scorecard and OpenAI's disclosures describe the same problem from two directions: containment keeps lagging capability, and most of what a customer can verify is whatever the vendor chooses to publish.
What Defenders Should Ask Their AI Vendors
Here is the practical value for a security team. If you are buying or deploying agentic AI, you are exposed to exactly the gap Guidelight measured, and you do not have to wait for a lab to volunteer its plan. Put the questions in the contract.
|
● Containment Questions for Your AI Vendor
Four things to get in writing before you deploy an agentic model, and the answer that should worry you.
|
|
Ask 1: The Written Containment Playbook
Request the rogue-model incident plan in the RFP: which permissions get revoked, who the model keeps serving, and when it is taken fully offline.
|
|
Ask 2: Kill Switch and Rollback
Confirm a tested way to pause workloads, restrict a model’s permissions, or shut it down, and ask when it was last exercised, not just whether it exists on paper.
|
|
Ask 3: Monitoring and Circuit Breaking
Ask whether internal AI activity is logged and scanned, and whether systems auto-halt after a surge of flagged misbehavior.
|
|
Ask 4: Independent Review and Disclosure
Ask who audits the controls, how deep their access goes, and whether the findings are published. METR participation is a reasonable floor.
|
|
and the red flag ↓
|
|
The Answer That Should Worry You
“We can’t share that.” A vendor may have legal reasons to stay vague, but an undisclosed containment plan and a missing one look identical from your side of the contract. Price the uncertainty into the deal.
|
|
Framework: Guidelight AI Standards Control standard, six control practices, August 2026. Checklist: The CyberSignal.
|
The four containment questions to put to any AI vendor, drawn from the six control practices Guidelight graded. Alt text: a stacked checklist diagram with four purple question cards over a red card reading the answer that should worry you, a vendor that will not share its containment plan.
Concretely, that means four things in your next vendor review:
- Request the written containment playbook in the RFP. Ask for the rogue-model incident plan in writing: what permissions get revoked, who the model keeps serving and under what limits, and the trigger for taking it fully offline. "We take safety seriously" is not an answer.
- Ask about the kill switch and rollback. Confirm there is a tested mechanism to pause workloads, restrict a model's permissions, or shut it down, and ask when it was last exercised.
- Ask about red-teaming and disclosure. Find out who independently reviews the vendor's controls, how much access those reviewers get, and whether findings are published.
- Treat "we can't share that" as a data point. A vendor may have sound legal reasons to be vague, but from your side of the contract, an undisclosed plan and a nonexistent one look the same. Price that uncertainty into the deal.
My Read
My read: the useful signal here is not the letter grades, it is the shape of the gap. Every lab in this study is better at noticing a problem than at stopping one, which is the wrong order if you believe these models are getting more capable and more autonomous. I would not read Meta's F or Anthropic's zero on containment as proof that either is reckless. Guidelight is grading a paper trail, and the labs' responses suggest more is happening privately than is written down. But that is cold comfort for a buyer, because "trust us, it is handled internally" is not a control you can audit, budget for, or fall back on when your own environment is the one an agent breaks out of. The report's real contribution is to convert a vague unease into six specific questions, and those questions work just as well pointed at a vendor as at a frontier lab.