Frontier AI Labs Still Have Major Gaps in Model Control

Monitoring is not prevention.
Lily Morris
Contributing Writer
electronic eye, technology, surveillance, security, monitoring concept
valerybrozhinsky - stock.adobe.com

Guidelight’s first assessment of frontier AI control practices finds that Anthropic, OpenAI, Google, xAI, and Meta have yet to fully implement basic safeguards for supervising their internal AI systems. The assessment covers logging, monitoring performance, controls that block high-risk actions, circuit breakers, independent review, and containment planning.

Anthropic and OpenAI lead with C+ grades and average scores of 2.50 out of 5. Google received a D+ (1.50), xAI a D− (0.83), and Meta an F (0.67).

No company scored above 3, or “substantial partial implementation,” on any practice.

Why It Matters: AI companies use their own models for sensitive work, including coding and cybersecurity. Guidelight finds that safeguards for observing model behavior are further along than systems for blocking dangerous actions or containing models after an incident. That gap becomes more consequential when agents receive greater access to internal systems and operate with less human supervision.

  • No Company Meets Guidelight’s Baseline in Full: Anthropic and OpenAI have the strongest controls in the assessment, especially in logging and monitoring. Google has implemented less, although Guidelight describes its AI Control Roadmap as the most specific published plan among the five companies. Meta and xAI have the weakest scores and less public evidence of deployed controls. Across all 30 company-practice scores, none earned a 4 or 5.
  • Monitoring Is Further Along Than Prevention: Anthropic and OpenAI each scored 3 for logging and monitor efficacy. Anthropic reports monitoring coverage for more than 90% of employees and has measured the recall of parts of its monitoring system. Gaps remain, including unmonitored inference and jobs that can run without logging. Google scored lower in these areas, while public evidence from Meta and xAI was more limited.
  • Controls That Stop Risky Actions Remain Incomplete: Guidelight wants sensitive changes to security and monitoring systems reviewed before execution, along with circuit breakers that restrict activity after repeated warnings. Anthropic’s Claude Code auto mode can block actions and stop or escalate an agent after repeated denials. OpenAI’s Codex auto-review can block certain actions that cross sandbox boundaries. Guidelight found no evidence, however, of a system-wide circuit breaker capable of restricting all relevant models following a serious incident.
  • Containment Is a Major Weakness, While Outside Review Is Further Along: Anthropic received the highest third-party review score, and most assessed companies participated in METR’s Frontier Risk Report. xAI did not. OpenAI received the highest containment score after documenting cases where it paused training or restricted model activity following safety incidents. Even there, Guidelight found no evidence of a formal standing containment plan for future misalignment incidents. Anthropic and Meta received no implementation credit for containment.
  • The Scores Depend on Public Evidence and Cover Only Part of AI Safety: Guidelight used information available through August 18, 2026 and gave companies an opportunity to provide additional public evidence or correct interpretations. Undisclosed safeguards could lead to lower scores than companies deserve, while unaudited claims could lead to higher ones. The assessment covers only part of Guidelight’s Control standard, and its current requirements do not cover model training or external testing. Guidelight plans to update the ratings as new information becomes available.

Go Deeper -> AI Control: An Assessment of Frontier Practices – Guidelight

Cybersecurity updates, executive insights, and the stories shaping the enterprise.

Browse past editions of TNCR newsletters. 

Technology news, cybersecurity, & executive insights.

×
You have free article(s) left this month courtesy of the CIO Professional Network.

Enter your username and password to access premium features.

Don’t have an account? Join the community.

Would You Like To Save Articles?

Enter your username and password to access premium features.

Don’t have an account? Join the community.

Thanks for subscribing!

We’re excited to have you on board. Stay tuned for the latest technology news delivered straight to your inbox.

Save My Spot For TNCR LIVE!

Thursday April 18th

9 AM Pacific / 11 PM Central / 12 PM Eastern

Register for Unlimited Access

Already a member?

Digital Monthly

$12.00/ month

Billed Monthly

Digital Annual

$10.00/ month

Billed Annually

Would You Like To Save Books?

Enter your username and password to access premium features.

Don’t have an account? Join the community.

Log In To Access Premium Features

Sign Up For A Free Account

Name
Newsletters