How AI Labs Approach Safety: Frontier Safety Frameworks Explained

Every few weeks, a major AI lab publishes a new model, a new safety policy or a long technical report. Each announcement uses its own vocabulary: “AI Safety Levels”, “Critical Capability Levels”, “High capability thresholds”, “system cards”. It is hard to tell what actually changed, and harder still to judge whether it matters.

This guide explains the common structure behind the safety programs of leading AI developers such as OpenAI, Anthropic, Google DeepMind and Meta. It covers what a frontier safety framework is, how the main frameworks compare, how to read a system card, what outside testing adds, and where these approaches fall short. It also explains how we put together our weekly AI Lab Safety Roundups.

Last reviewed: September 29, 2026.

What is a frontier safety framework?

A frontier safety framework is a public document in which an AI developer describes how it will identify and manage severe risks from its most capable models. Most frameworks follow an if-then logic:

  • If a model reaches a defined level of dangerous capability,
  • then the company commits to specific safeguards before training further or deploying it, and to pausing if it cannot put those safeguards in place.

The idea gained momentum in 2023 and became an industry expectation after the Seoul AI summit in May 2024, when 16 companies signed the Frontier AI Safety Commitments and agreed to publish such frameworks. Laws such as California’s SB 53 and the EU AI Act’s rules for general-purpose AI now require large developers to have and follow comparable safety documentation. We cover that legal context in our AI risk and policy guide.

The building blocks every framework shares

Risk domains

Frameworks focus on a short list of catastrophic or severe risks rather than every possible harm. Common domains include:

  • Chemical and biological weapons: whether a model could meaningfully help someone create or deploy them.
  • Cybersecurity: whether a model could significantly help discover and exploit vulnerabilities or automate cyberattacks.
  • Autonomy and AI research acceleration: whether a model could act independently in dangerous ways or dramatically speed up AI development itself.
  • Harmful manipulation and misalignment: whether a model could persuade or deceive at scale, or undermine human oversight.

Capability thresholds

Each framework defines levels at which a model’s capabilities are serious enough to require stronger protections. The names differ by company, but the concept is the same.

Evaluations

Labs run tests before and during development to estimate whether a model is near a threshold. These include automated benchmarks, expert-designed tasks, “uplift” studies comparing people with and without AI assistance, and red teaming.

Safeguards

Protections come in two families:

  • Deployment safeguards reduce misuse by users: refusal training, classifiers that block dangerous content, account-level monitoring, restricted access and usage policies.
  • Security safeguards protect the model itself, above all its weights, from theft by criminals or state-backed attackers, since a stolen model can be used without any safeguards.

Governance

Frameworks name who makes the call: an internal safety committee, senior leadership, the board, and in some cases outside advisors. They also describe how the framework itself can be changed.

How the major frameworks compare

The details change often, and each company revises its framework periodically, so always check the current version on the company’s own site. As of our last review:

DeveloperFrameworkThreshold vocabularyNotes
AnthropicResponsible Scaling Policy (first published 2023)AI Safety Levels (ASL-1, ASL-2, ASL-3 and above), loosely modeled on biosafety levelsActivated ASL-3 protections as a precaution when it launched Claude Opus 4 in May 2025; has revised the policy several times since.
OpenAIPreparedness Framework (first published 2023, major revision April 2025)“High” and “Critical” capability thresholds in tracked categories such as biology and chemistry, cybersecurity and AI self-improvementA Safety Advisory Group reviews results. In 2026 OpenAI said it would evolve the framework after finding an upcoming model may meet its Critical cybersecurity threshold.
Google DeepMindFrontier Safety Framework (introduced 2024, updated regularly)Critical Capability Levels (CCLs)Later versions added areas such as harmful manipulation and risks to human oversight.
MetaFrontier AI Framework (2025), updated and renamed the Advanced AI Scaling Framework in 2026Risk thresholds for cyber and chemical and biological risks; the 2026 version adds loss of controlCovers open-weight as well as API and closed releases, and introduces per-model Safety & Preparedness Reports.
OthersMicrosoft, Amazon, xAI and other signatories have published frameworks of their ownVariesDepth and specificity vary widely.

How to read a system card

A system card (or model card) is the document a lab publishes alongside a model to describe its capabilities, testing and safeguards. When a new one appears, we look for:

  1. Scope. Which model versions and products does it cover, and was testing done on the final release version?
  2. Dangerous-capability results. How did the model score against the framework’s thresholds, and what did the company conclude?
  3. Safeguards applied. Which deployment and security measures are in place, and are any new?
  4. Alignment and behavior testing. Results on honesty, sycophancy, refusals, jailbreak resistance, and any concerning behaviors observed in testing.
  5. External testing. Which outside organizations had access, how long they had, and whether their findings are summarized or published.
  6. Known limitations. What the company says it could not rule out or measure well.

A long system card is not automatically a good one. The most useful ones are specific about methods, candid about uncertainty and clear about which decisions were made on the basis of the results.

What external testing adds

Company self-assessment has obvious limits, so outside testing plays a growing role:

  • Government institutes such as the UK AI Security Institute and the US Center for AI Standards and Innovation have evaluated some frontier models before release under agreements with developers.
  • Independent evaluators such as METR and Apollo Research test models for autonomous capabilities and deceptive or scheming behavior.
  • Bug bounties and academic researchers probe safeguards such as jailbreak defenses after release.

External testing is only as strong as the access testers get. Time, model versions, and whether results can be published all shape what it can tell us.

The limits of self-governance

Frontier safety frameworks are a real step forward. They make commitments explicit and create a public record to hold companies to. Critics point to several weaknesses:

  • Companies write, interpret and revise their own rules. A framework can be updated when it becomes inconvenient.
  • Thresholds can be vague. Terms like “significant uplift” leave room for judgment, and evaluations may underestimate what a determined user could achieve.
  • Commercial pressure. Competition encourages faster releases, which can compress testing time.
  • Coverage gaps. Frameworks focus on catastrophic risks and say less about everyday harms such as bias, privacy and misinformation.
  • Open weights. Once model weights are published, deployment safeguards cannot be enforced, which changes the risk calculation.

Regulation is starting to address some of these gaps by requiring frameworks to exist, be followed and be reported on, but enforcement is still new.

How our weekly roundup works

Each week we review official announcements, system cards, policy updates, research papers and credible reporting on the safety work of major AI developers. For every item, we try to answer four questions:

  1. What was released or announced?
  2. What does the company say it tested, and how?
  3. Has anyone outside the company verified it?
  4. What remains unknown?

We link to primary sources, clearly separate company claims from independent findings, and avoid treating announcements as proof of safety. You can see the format in action in our roundup of OpenAI’s assessments and Anthropic’s new commitments. Our standards are described in our editorial policy.

Why this matters beyond the labs

Decisions made inside a handful of companies shape the tools millions of people use every day. Understanding frameworks and system cards helps you judge claims about new models, compare providers and follow the policy debate. For the fundamentals of why these safeguards are needed, see AI safety and alignment explained. For how safety shows up in the products you use, see our AI tool safety reviews, and if you are building with these models, see responsible AI for developers.

Frequently asked questions

What is Anthropic’s Responsible Scaling Policy?

It is Anthropic’s frontier safety framework. It defines AI Safety Levels (ASLs) based on a model’s potential for catastrophic misuse or autonomy and specifies the security and deployment safeguards required at each level before a model can be trained further or released.

What is OpenAI’s Preparedness Framework?

It is OpenAI’s process for tracking and managing severe risks from frontier models. It defines capability thresholds in categories such as biology and chemistry, cybersecurity and AI self-improvement, and requires safeguards to be in place before models reaching those thresholds are deployed.

What is a system card in AI?

A system card is a document published with an AI model that describes its capabilities, the safety evaluations it underwent, the results, the safeguards in place and its known limitations.

Are AI lab safety frameworks legally binding?

The frameworks themselves are company policies, but laws increasingly require them. California’s SB 53, for example, requires large frontier developers to publish and follow a safety framework, and the EU AI Act sets obligations for providers of general-purpose AI models with systemic risk.

Which AI lab is the safest?

There is no agreed ranking. Labs differ in how specific their frameworks are, how much external testing they allow, how transparent they are about results, and how they handle open-weight releases. We compare these factors in our weekly roundups rather than issuing a single score.

Sources

Primary sources used for this guide, checked on September 29, 2026. Policies and products change, so always check the latest version at the source.

Latest AI lab safety roundups

John

Written by

John

Ellis Marlow writes Safely Clever's guides on AI safety, AI tool safety reviews, responsible AI for developers, AI risk and policy, and AI lab safety. Ellis tests AI tools hands-on, reads the safety research, system cards and lab announcements, and explains them in plain English, with links to the primary sources. Articles are reviewed by the Safely Clever editorial team and updated when the facts change. Ellis Marlow is a pen name.