AI Safety Frameworks: What the Labs Actually Promised
The labs wrote the rules, set the thresholds and grade their own homework — and one added a clause letting it lower the bar if a rival ships first.
For most of the AI boom, “we take safety seriously” was a line in a blog post, not a document you could hold a company to. That has changed, at least on paper. The three largest Western labs — Anthropic, OpenAI and Google DeepMind — have each published a formal safety framework: a written policy that says, in effect, here are the specific dangerous things our models might one day be able to do, here is how we will test for them, and here is the point at which we will hold a model back rather than release it.
This is genuinely new, and worth taking seriously. It is also, on inspection, a set of rules the referees wrote for themselves, grade themselves against, and can rewrite when the season gets difficult. Both of those things are true at once, and the interesting part of the story lives in the gap between them.
We have written before about what the phrase “AI safety” is actually doing when a company uses it. This piece is narrower and more concrete: the frameworks themselves — what they commit their authors to, what they conspicuously leave out, and how much weight the word “commitment” can bear when nobody outside the company can enforce it.
What a “safety framework” actually is
Strip away the branding and every frontier safety framework has the same four moving parts. First, a set of capability thresholds: specific dangerous things a model might be able to do — meaningfully help a novice build a bioweapon, run an autonomous cyber-attack, or sharply accelerate AI research itself. Second, an evaluation regime: tests run on new models to check how close they are getting to those thresholds. Third, a set of mitigations that must be in place before a model at a given level ships — security to stop the weights being stolen, and deployment safeguards to stop the capability being misused. Fourth, and most importantly, a precommitment: a promise that if the mitigations cannot hold the risk below the threshold, the model does not ship.
That last part is the genuinely novel bit. A company saying “we will not release this product, even though we could, because it is too dangerous” is not something the technology industry has historically done voluntarily. Whether any of them would actually do it under commercial pressure is exactly the open question. But the pledge is now written down, which is more than was true two years ago.
The big three, and what each one commits to
The frameworks share a skeleton but differ in the detail and the vocabulary:
- Anthropic’s Responsible Scaling Policy (first published September 2023, now in its third version) organises risk into AI Safety Levels, or ASLs, explicitly modelled on the biosafety levels used for laboratory pathogens. Higher ASLs trigger stricter security and deployment requirements. In May 2025 Anthropic activated the ASL-3 standard for Claude Opus 4 — the first time a lab has publicly stepped up its safeguards under its own framework — calling it a precautionary move because it could not rule out that the model crossed the line.
- OpenAI’s Preparedness Framework (first published December 2023, substantially updated in April 2025) tracks capabilities in categories such as biological, chemical and nuclear uplift, cybersecurity, and AI self-improvement, with “High” and “Critical” thresholds that gate deployment and further development.
- Google DeepMind’s Frontier Safety Framework (first published May 2024, updated to version 2.0 in February 2025 and strengthened again since) uses Critical Capability Levels across CBRN, cyber, machine-learning R&D and, notably, deceptive alignment — the risk of a model actively working to undermine human oversight.
The convergence is not an accident. It is partly the result of the same small pool of technical-safety researchers moving between labs, and partly the product of an international process that asked every major company to produce one of these documents to a common shape.
The part worth crediting
It would be lazy to wave all this away as public relations, so let us give it its due. Before these frameworks, the honest answer to “what would make you not release a model?” was “trust us.” Now there is at least a written threshold to point at, and a couple of concrete things have actually happened because of it. Anthropic really did turn on heavier safeguards for Opus 4. The labs now submit frontier models to external evaluators — groups such as METR and Apollo Research — for dangerous-capability and deception testing before release, a practice that barely existed in 2022.
There is also a template effect. A written framework is something a regulator, an auditor or a court can eventually get hold of, compare against behaviour, and use as a yardstick. It is far easier to hold a company to a threshold it published than to a mood it once radiated on stage. The frameworks are, at minimum, a floor to argue from — and a floor is more than the industry had.
Who wrote the rules? The people they bind
Now the other half. Every one of these frameworks is written by the company it governs, sets thresholds that company chose, is tested by evaluations that company designed, and is graded by that company’s own assessment of whether it complied. There is no external body with the authority to say “your evaluation was inadequate, this model does cross the line, you may not ship it.” The safety case and the launch plan are produced by the same organisation, under the same commercial pressure, to the same deadline.
Independent scorecards that try to grade the labs from outside are not flattering. The Future of Life Institute’s 2025 AI Safety Index, which borrows a risk-management taxonomy from the non-profit SaferAI, gave no company a grade better than C+, and none better than D on planning for the human-level systems several of them say they are explicitly trying to build. One reviewer flagged the absence of any “coherent, actionable plan” for controlling such systems as deeply worrying. When the people building the thing grade their own safety homework, the marks are middling; when outsiders grade it, the marks are worse.
The escape hatches
A commitment matters only if it holds when keeping it is expensive. This is where the frameworks get slippery. They are, by design, living documents — which means they can be revised, and the revisions have not all pointed towards caution. In April 2025 OpenAI updated its Preparedness Framework to add that if a competitor released a high-risk system without comparable safeguards, it “may adjust” its own requirements. The reasoning is candid and, on its own terms, rational: no company wants to unilaterally hold back while a rival ships. The effect is a written permission slip to lower the bar precisely when the competitive race is at its most dangerous.
The softening language is everywhere once you look for it. Anthropic activated ASL-3 as a “precautionary” and “provisional” measure, having not actually determined the model required it. Thresholds are hedged with words like “appropriate” and “where feasible.” Publication deadlines slip. None of this is necessarily bad faith — a genuinely uncertain science needs room to update — but it does mean the brake and the accelerator are wired to the same pedal, controlled by the same foot.
The international scaffolding around the frameworks
The company frameworks did not appear in a vacuum; they sit inside a loose international structure assembled at a run of summits. In November 2023, twenty-eight countries and the EU signed the Bletchley Declaration, the first multilateral statement to acknowledge frontier-AI risk. At the Seoul summit in May 2024, sixteen companies signed the Frontier AI Safety Commitments — the pledge that produced most of the frameworks above — agreeing to publish thresholds and, in the strongest line, to “not develop or deploy a model or system at all” if risks could not be kept below them. The signatory list has since grown past twenty.
Two developments are worth noting. The first is the International AI Safety Report, chaired by Yoshua Bengio and written by ninety-six experts nominated by thirty countries plus the UN, EU and OECD, whose first full edition landed in January 2025 — a serious attempt at a shared scientific baseline rather than a marketing document. The second is the mood shift: at the Paris AI Action Summit in February 2025, the framing moved from “safety” to “action” and opportunity, and the United States and the United Kingdom declined to sign the closing declaration at all. The scaffolding exists; the political will holding it up is visibly wobbling. This is the layer where voluntary pledges would, in a firmer world, harden into the kind of enforceable regulation we have looked at elsewhere.
What the frameworks quietly leave out
Read the thresholds closely and you notice what they are about: catastrophe. Bioweapons, cyber-attacks, autonomous replication, loss of human control. These are the right things to worry about at the extreme, and also the things least likely to touch you this year. The frameworks are largely silent on the harms people actually meet — the confident wrong answer, the model that agrees with whatever you say, the biased screening decision, the quiet data leak. Those are filed as product-quality issues, not “safety,” and so fall outside the one document that carries the word.
There is a measurement problem underneath this too. A 2024 paper introduced the term safetywashing for a specific finding: many benchmarks that claim to measure “safety” are, statistically, mostly measuring general capability — they climb as models get bigger and more capable, whether or not anything actually got safer. If the tests that populate a framework’s evaluations cannot cleanly separate “more capable” from “more safe,” a lab can report improving safety scores while shipping a model that is only more capable across the board, dangerous parts included.
So are they worth anything?
Yes, but not the amount the word “commitment” implies. Treat a safety framework as the best current statement of what a lab believes it should do, published in a form that can be checked against its behaviour later. That is real value: it creates a paper trail, forces internal argument before a launch, gives external evaluators a door in, and hands regulators a ready-made template. It is a floor.
What it is not is a guarantee, because the missing ingredient is enforcement. A promise you write, mark and can revise yourself is a policy, not a constraint, and it is weakest at exactly the moment — a heated competitive race — when a constraint would matter most. Until an outside body can audit the evaluations and veto a launch, the frameworks remain an honour system run by the least disinterested party. That gap is also why the question of who is actually liable when an AI system causes harm keeps landing back on courts and regulators rather than on the frameworks themselves.
How to read a safety framework without being sold one
You do not need to parse forty pages of policy to take the measure of one. A few questions do most of the work:
- Who checks the homework? Are the evaluations run or verified by anyone outside the company, or graded entirely in-house?
- What is the escape clause? Look for the conditions under which the company may lower its own bar — competitive pressure, “provisional” activations, feasibility caveats.
- Has it ever cost them anything? A framework that has never delayed, restricted or blocked a release is a framework that has never actually bound.
- Does it cover the harms you will meet? Catastrophic-risk thresholds matter, but they are not the same as your day-to-day exposure to unreliability, bias and lost privacy.
The frameworks are a real step, and a better world than the one where “trust us” was the entire policy. But a rule is only as good as the person who can enforce it, and right now that person is the same one holding the release schedule. Read them as promises made in good faith and hedged in self-interest — useful, watchable, and not yet the thing that will actually stop a dangerous model from shipping on time.
Frequently asked questions
Are AI safety frameworks legally binding?
No. Anthropic’s Responsible Scaling Policy, OpenAI’s Preparedness Framework and Google DeepMind’s Frontier Safety Framework are voluntary self-governance documents, not law. Each company writes, interprets and enforces its own. The nearest hard law is the EU AI Act’s rules for general-purpose models, which impose obligations but do not replace the frameworks. Outside the EU, there is currently no external body that can audit a lab’s safety evaluation or block a launch that breaches its own stated threshold.
What is the difference between Anthropic’s ASLs and OpenAI’s Preparedness levels?
Both are capability-threshold systems, but they use different vocabulary. Anthropic’s AI Safety Levels (ASLs) are modelled on the biosafety levels used for laboratory pathogens: higher ASLs trigger stricter security and deployment requirements. OpenAI’s Preparedness Framework instead tracks specific risk categories — biological, chemical and nuclear uplift, cybersecurity, AI self-improvement — with ‘High’ and ‘Critical’ thresholds that gate deployment. Google DeepMind’s framework uses ‘Critical Capability Levels’ across a similar set of domains plus deceptive alignment.
Has any lab ever actually held back a model because of its framework?
The clearest concrete action so far is Anthropic activating its ASL-3 standard for Claude Opus 4 in May 2025 — heavier security and deployment safeguards — which it described as a precautionary, provisional step because it could not rule out that the model reached the threshold. That is a real, costly action taken under a framework. No lab has publicly cancelled a major model outright on safety grounds, so the strongest version of the promise — refusing to ship at all — remains untested.
Can a company change its safety framework whenever it wants?
Yes. All three are explicitly living documents that can be revised, and not every revision has tightened the rules. In April 2025 OpenAI updated its Preparedness Framework to add that if a competitor released a high-risk system without comparable safeguards, it ‘may adjust’ its own requirements. The logic is candid — no company wants to hold back alone while a rival ships — but the effect is a written route to lowering the bar exactly when competition is fiercest.
Do these frameworks cover everyday harms like bias or hallucination?
Mostly no. The thresholds target catastrophic misuse — bioweapons, large-scale cyber-attacks, autonomous replication, loss of human control. The harms most people actually meet — confident wrong answers, sycophancy, biased screening, quiet data leaks — are generally treated as product-quality issues, not ‘safety’, and fall outside the frameworks. So a model can be fully compliant with its safety framework and still be unreliable, biased or intrusive in ordinary use.
Sources
- Anthropic’s Responsible Scaling Policy — Anthropic
- Activating AI Safety Level 3 Protections — Anthropic
- Updating our Preparedness Framework — OpenAI
- Strengthening our Frontier Safety Framework — Google DeepMind
- Frontier AI Safety Commitments, AI Seoul Summit 2024 — GOV.UK
- The Bletchley Declaration by countries attending the AI Safety Summit, 1–2 November 2023 — GOV.UK
- International AI Safety Report 2025 — International AI Safety Report
- AI Risk Management Framework (AI RMF 1.0) — NIST
- AI Safety Index — Summer 2025 — Future of Life Institute
- Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress? — NeurIPS 2024
