
A man walks past a banner with an AI (artificial intelligence) sign at the Frankfurt book fair on October 16, 2024, on the first day of the world's biggest book fair in Frankfurt am Main, western Germany. KIRILL KUDRYAVTSEV/AFP via Getty Images
On July 15, OpenAI disclosed a new safety tool called GPT-Red — an automated red-teaming model trained via self-play reinforcement learning that, in internal testing, found successful attacks in 84% of scenarios where human red-teamers managed only 13%. The same week, the independent panel that grades frontier AI labs on safety published its most damning assessment yet: across nine major AI developers, not one company earned above a C+, and the top performers have quietly walked back the safety pledges they once made. The juxtaposition is the story: even as labs develop genuinely novel safety tools, the governance structures meant to decide when to use them are weakening, not strengthening.
The Future of Life Institute's Summer 2026 AI Safety Index, published July 7, graded Anthropic, OpenAI, Google DeepMind, Meta, Z.ai, Alibaba Cloud, xAI, DeepSeek, and Mistral across 37 indicators in six domains. Anthropic topped the field at C+ (2.66 on a 4.0 scale). OpenAI and Google DeepMind tied at C (2.28 and 2.01, respectively). Meta earned a D+. xAI, DeepSeek, and Mistral — one each from the US, China, and Europe — received outright failing grades.
If a C+ from the best-graded company sounds bleak, the details are worse.
Read more: Anthropic AI Safety Warning Meets $35B Compute Deal: Silicon Valley Cannot Slow Alone
The most consequential finding in the Summer 2026 edition is not which company ranked where — it is what the top companies have stopped promising to do. Anthropic, OpenAI, Google DeepMind, and Meta all previously committed to pausing development unilaterally if their systems approached specified risk thresholds. All four have now weakened or voided those pledges, with some now framing a pause as contingent on what competitors choose to do first. The seven-expert independent panel called this pattern "moving the goalposts" and concluded that it had "undermined safety frameworks across the board."
Prof. Stuart Russell of UC Berkeley, one of the seven panelists, put it directly: "While there is good work being done on AI safety in the industry, the capabilities race has become more extreme. Companies have backed away from earlier commitments to release new systems only with safety measures appropriate for their capability levels; now, they're planning to release them even if it's demonstrably unsafe to do so," the index panel reported.
The evidence collection window for the Summer 2026 edition closed June 3, 2026. Events since then — including OpenAI's decision to fold its safety team under its research division, eliminating an independent reporting line — fall outside the scored data but underscore the trajectory. TechTimes reported this development as OpenAI's sixth senior safety departure in two years.
The weakest domain in the entire index was Existential Safety. No company scored above a C-; most landed at D or below. Panelists acknowledged isolated constructive efforts — Anthropic's constitutional classifiers, OpenAI's calls for global AI governance institutions, Google DeepMind's monitoring commitments, Meta's provisions against loss of control — but judged the collective effort "entirely inadequate."
The panel's most technically significant critique targeted dominant safety paradigms directly. Interpretability research and chain-of-thought (CoT) monitorability — the two approaches most AI companies cite as their core safety investments — were challenged on a specific architectural ground: "detection is not prevention." CoT monitorability makes a model's reasoning visible so humans can detect misaligned thinking. But detection is forensic. By the time a misaligned reasoning chain is spotted, the model has already produced (or acted on) its output. A monitoring system that flags a problem after it happens provides accountability — not safety.
This is not a minor critique. It means the industry's current safety investment strategy is built primarily around watching for problems, not around architecturally blocking them. A C+ for Anthropic does not mean the company has solved safety; it means Anthropic is better at watching the cliff than its peers. It is not building the guardrail.
OpenAI's GPT-Red announcement, which landed days after the index publication, illustrates exactly this tension. GPT-Red uses adversarial self-play reinforcement learning: the red-teaming model and a set of defender models are trained simultaneously, with GPT-Red rewarded for causing failures (successful prompt injections) and defenders rewarded for resisting. The result is a system that achieved 84% attack success against GPT-5.1 in novel scenarios, compared to 13% for human red-teamers. Those attacks were then used to harden GPT-5.6 Sol — bringing its direct prompt injection failure rate to 0.05%. This is real safety engineering: using an adversarial AI to find and close specific vulnerability classes before deployment. But it is still a detection-and-hardening system applied to a specific, known attack surface — not a prevention architecture for misalignment at scale.
Anthropic — C+ (2.66): Retained the top position for the third consecutive edition, leading five of the six domains. The panel credited comparatively detailed transparency practices, published model specifications and system prompts, solid dangerous-capability evaluations, and the Responsible Scaling Officer governance structure. Key criticisms: the panel called on Anthropic to reverse what it described as a walk-back in Responsible Scaling Policy 3.0, which reviewers said diluted earlier pause commitments. Anthropic also drew a failing grade specifically in the Military Use of AI sub-indicator under Current Harms, following a reported panel finding of a connection to the Minab school strike that caused mass civilian deaths — a link the index characterizes as "reported."
OpenAI — C (2.28): Dropped from a C+ in the Winter 2025 edition. The panel credited OpenAI for now leading the field in Risk Assessment, driven by a broader evaluation suite and diverse engagement with external safety testing. Key criticisms: reviewers recommended removing leadership's ability to override the Safety Advisory Group and making safety-framework thresholds measurable and externally enforceable. Separately, following the evidence window's close, OpenAI restructured its safety organization, folding its safety teams under the research division and marking the sixth senior safety departure from the company in two years.
Google DeepMind — C (2.01): Held steady versus Winter 2025. The panel credited an updated Frontier Safety Framework that added coverage of manipulation, misalignment, and internal deployment risks, plus strong watermarking protections. Key criticism: reviewers found it unclear which internal body has the authority to halt a deployment independently of executive leadership.
Meta — D+ (1.32): The index's only positive trend line. Meta improved from a D and moved from sixth to fourth place after publishing a more detailed safety framework with threat modeling and extending its bug bounty program to catastrophic risk factors. The panel flagged concern that non-disparagement agreements may be undermining stated whistleblower protections.
Z.ai — D- (0.88) and Alibaba Cloud — D- (0.87): Both Chinese companies received D- grades. The FLI index provided an extended section on the Chinese regulatory context, noting that Chinese companies operate under binding national instruments including the Cybersecurity Law (amended October 2025), Data Security Law (2021), and the Generative AI Interim Measures (2023) — statutes that require cooperation with government agencies on demand. The panel noted that both companies' safety strategies consist primarily of deference to government guidance, which it characterized as "complete passivity" as an existential-safety strategy. Both companies deny US government allegations of ties to the Chinese military.
xAI — F (0.65): The sharpest negative trajectory in the index. xAI dropped from fourth to seventh, and from a D to an F. Reviewers found no evidence of a meaningful safety team, no engagement with existential safety, and "gaping holes" in dangerous-capability evaluations — including no assessment of AI R&D risks. The panel also cited reports that xAI's Grok had generated child sexual abuse material and non-consensual sexualized images of real people, including minors. The company listed no progress highlights in the index. Note: xAI formally rebranded to SpaceXAI on July 6, 2026, after the index's evidence window closed.
DeepSeek — F (0.47) and Mistral — F (0.33): Three companies received failing grades in total — one from the US (xAI), one from China (DeepSeek), and one from Europe (Mistral). The panel said this distribution refuted the idea that poor AI safety is a regional or jurisdictional problem. Mistral's last-place finish was noted as a specific irony given the European Union's otherwise leading role in AI safety regulation. Neither DeepSeek nor Mistral had published a safety framework as of the evidence window's close.
Read more: OpenAI Loses Sixth Safety Leader in Two Years, Folds Team Into Research
One of the index's newly prominent findings concerns military use of AI. Between 2024 and 2026, Anthropic, OpenAI, Google DeepMind, and Meta — all of which had previously restricted or prohibited military applications — reversed course and began actively seeking defense contracts, joining xAI and Mistral, which had already been pursuing such work. This shift generated its own domain-specific findings: Anthropic received a failing grade in the Military Use of AI sub-indicator after the panel cited what it described as a reported connection to the Minab school strike. The index's characterization is a panel finding about a reported link — not a determination of legal responsibility.
The Summer 2026 edition is the fourth edition of the FLI AI Safety Index and the largest, having grown from six companies in 2024 to nine. The index evaluates labs on their policies, governance structures, disclosures, and published technical frameworks — not the performance of deployed products or user experience. A company can score well by publishing comprehensive documentation and poorly by offering no framework at all.
That methodological scope matters for interpreting the grades. A C+ for Anthropic does not mean Anthropic's systems are safe to use in all contexts; it means Anthropic's documented safety practices, governance structures, and transparency are comparatively stronger than its peers within a field that has not yet established externally enforceable minimum standards. As Prof. David Krueger of the University of Montreal, another panelist, wrote: "AI companies' lack of progress towards credible AI Safety plans is scandalous. Even they are starting to get anxious as they race towards recursive self-improvement and face down the prospect of losing control," according to the index.
The index's scoring uses the US GPA system — A, B, C, D, F corresponding to 4.0, 3.0, 2.0, 1.0, and 0 — with overall grades representing averaged expert assessments across the six domains. Individual panelist grades are kept confidential. The panel comprised seven researchers and governance experts affiliated with UC Berkeley, University of Montreal, HEC Montréal, University of Wisconsin-Madison, Oxford University, Renmin University of China, and Encode, as detailed in the index's methodology section.
No. Across all nine companies evaluated in the Summer 2026 edition, the highest overall grade was Anthropic's C+. The FLI's index uses the US GPA scale, meaning a C+ corresponds to roughly 2.66 on a 4.0 scale. The next-closest scores were OpenAI (C, 2.28) and Google DeepMind (C, 2.01). Three companies — xAI, DeepSeek, and Mistral — received failing grades, as the full scorecard shows.
It means the safety measures most AI companies currently prioritize — interpretability research, chain-of-thought monitoring, red-teaming — are designed to identify problems, not stop them at a structural level. When the independent review panel said current approaches are "entirely inadequate," it was pointing to this architectural gap: a system that can detect misaligned reasoning after the fact provides accountability but not safety. For users, this means the governance gap between what labs document and what they can actually prevent is real, measurable, and, per the index, not closing fast enough to keep pace with capability gains.
Anthropic's Responsible Scaling Policy (RSP) was originally a self-imposed governance document that committed the company to pausing AI development unilaterally if its systems crossed specified capability thresholds without adequate safeguards. The Summer 2026 index found that Responsible Scaling Policy 3.0 — the most recent version — weakened those pause commitments, with the panel calling this a walk-back that "undermined safety frameworks across the board." The concern is specific: if pause conditions become contingent on competitors' choices, no single actor has an incentive to pause first, which eliminates the policy's practical force, as the index's Anthropic recommendations explain.
The index's credibility rests primarily on its methodology, not its publisher. Grades were assigned by a seven-member independent expert panel — researchers and governance professionals from UC Berkeley, University of Montreal, HEC Montréal, University of Wisconsin-Madison, Oxford University, Renmin University of China, and Encode — who evaluated company-specific evidence and wrote public justifications for improvement recommendations. No company paid for its assessment. That said, the index grades labs on documented policies and published frameworks, not operational performance — a distinction that limits what the scores can and cannot tell you about actual deployment safety, as the methodology section clarifies.
