Guidelight grades frontier AI labs on rogue-model containment and control
Guidelight AI Standards published its first public Control assessment (info current through Aug 18; TechCrunch coverage Aug 22) of Anthropic, Google, Meta, OpenAI, and xAI across six practices: logging, monitor efficacy, gated actions, circuit breaking, third-party review, and containment plans. No lab scored above 3/5 on any practice; Anthropic and OpenAI tied at C+ (2.50), Google D+ (1.50), xAI D− (0.83), Meta F (0.67). Containment plans scored weakest overall—OpenAI led that category (3) while Anthropic and Meta scored 0 on public evidence. Distinct from Anthropic’s Aug Risk Report and from OpenAI’s pacing/cyber preparedness framework posts.





