AI Safety: How to Read a Lab's Policy and Tell If It Commits to Anything
In May 2025, Anthropic formally activated a higher tier of protections for one of its models. It was not a statement of good intentions: it came with content classifiers and more than a hundred security controls behind it. Telling a document that binds from one that merely declares is a concrete skill, and three signals settle it.
Every major AI lab now publishes a safety document. They all sound responsible. And they differ enormously: some bind their author to do something checkable at a specific moment, and others describe an intention.
The difference is not rhetorical, and you can spot it without being an expert. Here is the method, with real documents on the table.
What a binding commitment looks like
On 22 May 2025, Anthropic formally activated ASL-3 protections for its Claude Opus 4 model. ASL-3 is a level defined in advance in its Responsible Scaling Policy, and activating it is not a declaration: it triggers concrete measures.
The ones the announcement itself describes: on the deployment side, «constitutional classifiers» that, in its words, «monitor model inputs and outputs and intervene to block a narrow class of harmful CBRN information» — chemical, biological, radiological and nuclear weapons. On the internal security side, more than a hundred controls, including egress bandwidth limits to make extracting the model weights harder, because the standard «involves increased internal security measures that make it harder to steal model weights».
Note the structure of that commitment, because it is what to look for in any other: a threshold defined before it is reached, a verifiable moment when it is declared reached, and a list of measures that activate as a result. You can check whether it activated. You can check when. You can argue about whether the measures suffice.
The pattern repeats, and that is the interesting part
Google DeepMind published its Frontier Safety Framework on 17 May 2024, described as «a set of protocols for proactively identifying future AI capabilities that could cause severe harm and putting in place mechanisms to detect and mitigate them».
Its centrepiece is the critical capability levels: the minimum capability levels a model would have to reach to play a role in severe harm. It is the same architecture as Anthropic's under another name — define the threshold that triggers the response in advance, rather than improvising when it arrives.
That two competitors arrived separately at the same shape says something: it is the only shape that lets compliance be checked from outside. A policy without a threshold cannot be broken, which is why it cannot be kept either.
Where all of this stops
Now the part to be clear about before feeling reassured.
At the 2024 Seoul summit, twenty companies — among them Amazon, Anthropic, Google, IBM, Meta, Microsoft, Mistral AI, OpenAI, Samsung, xAI and NVIDIA — signed the Frontier AI Safety Commitments. The text is explicit about its own nature: signatories undertake to develop and deploy their frontier models responsibly «in accordance with the following voluntary commitments».
Voluntary. That word is in the document, not supplied by us, and it is exactly what to retain: they include setting intolerable risk thresholds and not deploying if those are crossed, but whoever defines the threshold, whoever measures, and whoever decides it has been crossed is, ultimately, the same company. No third party certifies.
The exception is the regulatory route. Regulation (EU) 2024/1689 imposes enforceable obligations on general-purpose models deemed to carry systemic risk — evaluation, serious-incident reporting, cybersecurity — with penalties behind them. There, non-compliance has consequences that do not depend on the goodwill of the party bound. The chapter dealing with these models is worth opening: it reads more easily than its reputation suggests.
What those thresholds name, and why it is the most informative part
One detail tends to get skipped, and for an attentive reader it is the richest part of these documents: in defining what they are guarding against, labs reveal what they consider plausible.
Anthropic does not describe its classifiers as a general filter for unpleasant content. It scopes them to «a narrow class» of information about chemical, biological, radiological and nuclear weapons. That precision cuts both ways: it says what worries them enough to warrant a dedicated classifier, and it says what falls outside that particular measure.
DeepMind, for its part, anchors its critical levels in the capability to «play a role in severe harm» within high-risk domains. Again: not a promise that the model will behave well. A claim about capability thresholds in specific areas.
And there is a second revelation, at the other end of the document. When Anthropic spells out that its internal security measures aim to make stealing the model weights harder, it is saying something no marketing campaign would: that it considers there are actors with both the interest and the means to take them. The security section of a policy like this is usually more informative about the real state of the industry than any interview.
So it is worth reading these texts hunting for concrete nouns — which capability, in which domain, against whom — rather than adjectives. Adjectives are free; nouns commit.
The three signals, checkable in ten minutes
1. Is there a threshold defined BEFOREHAND, or only a principle? «We will rigorously evaluate risks» is not a commitment: it is a description of intent. «When a model reaches this specific capability, we will activate these specific measures» is one. Look for the conditional with content in it.
2. Has there been an actual activation, with a date? A framework published and never triggered can mean two things: that no model crossed the threshold, or that the threshold sits where it does not get in the way. Documented activations — with their date and their model — are the best evidence the mechanism is not decorative.
3. Who verifies? This is where almost everything falls down. If whoever measures, whoever sets the bar and whoever decides it was crossed are the same organization, what you have is a self-assessment. It may be honest and serious; it remains a self-assessment. External verification — a regulator, a public institute, an independent audit — is a different category.
Why this is not a specialist's concern
It might seem these policies matter only to people working in the field. It is the other way round, for a practical reason: these documents are, as things stand, the main public source of information about what the manufacturers themselves know about their systems' limits.
When a lab defines a dangerous-capability threshold, it is saying what it considers dangerous and how far away it believes it is. When it activates a protection level, it is conceding that its model reached a capability it did not have before. That is information no product press release provides, written by the people who know most about it.
Reading them with judgment is not surveillance: it is the best available window into what is happening inside.
The deep end, undiluted
The four documents behind this piece are public and can be read without intermediaries:
- Anthropic (22 May 2025), the ASL-3 activation, with the list of measures it triggers.
- Google DeepMind (17 May 2024), the Frontier Safety Framework and its critical capability levels.
- Frontier AI Safety Commitments from the Seoul summit, with the full list of the twenty signatories.
- Regulation (EU) 2024/1689 on EUR-Lex, with a language selector, for the route that is actually enforceable.
The capability you leave with: faced with any safety policy, look for the threshold defined in advance, check whether it has ever been activated with a date, and ask who verifies. Three questions, ten minutes, and the difference between a commitment and a brochure stops being a matter of trust.
This article was produced with artificial intelligence under human editorial oversight.