IA 360
Current Affairs

Anthropic launches Claude Opus 4 with ASL-3 safeguards

Anthropic introduces Claude Opus 4 and Sonnet 4 and provisionally activates ASL-3 measures for Opus. The company had not confirmed that the model crossed the threshold requiring them.

6 min read AI-generated Leer en español
Anthropic launches Claude Opus 4 with ASL-3 safeguards

On May 22, 2025, Anthropic introduced Claude Opus 4 and Claude Sonnet 4 and activated ASL-3 measures for Opus. The original source supports the documentary core of the event; activation was precautionary and provisional; it did not confirm that the model had crossed the threshold. Terms not published by the party are attributed to the source that documented them.

The Models: Anthropic's Best Coder Yet

Anthropic is calling Claude Opus 4 "the best coding model in the world," posting 72.5% on the SWE-bench benchmark and 43.2% on Terminal-bench. What sets it apart isn't just the score but its stamina: the company says it can work continuously for several hours on complex, thousand-step tasks, a clear step up from previous Sonnet models. Rakuten, one of the partners cited in the announcement, put that claim to the test with an open-source refactor that the model ran autonomously for seven hours. Document supporting the figure.

Sonnet 4, the update to Sonnet 3.7, scores 72.7% on SWE-bench — slightly ahead of Opus 4 on that particular benchmark, though Anthropic notes it doesn't match its bigger sibling across most domains. Its goal is different: offering the best mix of capability and practicality for everyday use. GitHub has already announced that Sonnet 4 will power the new coding agent in GitHub Copilot. Document supporting the figure.

Both models are hybrid: they respond almost instantly or switch into an "extended thinking" mode for deeper reasoning, and can now use tools — like web search — during that reasoning process, alternating between thinking and acting. Anthropic also notes that both models are 65% less likely than Sonnet 3.7 to take shortcuts or exploit loopholes to complete tasks artificially, a known issue with agentic models. Document supporting the figure.

Claude's Pro, Max, Team, and Enterprise plans include both models with extended thinking, and Sonnet 4 is also available to free users. Pricing holds steady from previous generations: Opus 4 costs $15 and $75 per million input and output tokens respectively, while Sonnet 4 costs $3 and $15. Both are available via API, Amazon Bedrock, and Google Cloud's Vertex AI. Document supporting the figure.

Claude Code Moves Out of Preview

Alongside the models, Anthropic confirmed the general availability of Claude Code, its AI-assisted coding tool, following a preview period the company describes as having received extensive positive feedback. Claude Code now supports background tasks via GitHub Actions and native integrations with VS Code and JetBrains, displaying edits directly in the developer's files. Anthropic is also releasing a Claude Code SDK so third parties can build their own agents on the same foundation, plus four new API capabilities: the code execution tool, an MCP connector, a Files API, and the ability to cache prompts for up to one hour.

Why Opus 4 Is Deployed Under ASL-3

The most consequential part of the announcement isn't about performance — it's about safety. Anthropic has simultaneously activated the ASL-3 standards under its Responsible Scaling Policy (RSP), a framework that defines increasing levels of protection based on a model's estimated danger. Until now, all of the company's models operated under ASL-2, which includes training the model to refuse dangerous CBRN requests and basic defenses against the theft of its weights (the parameters that make up its intelligence).

Anthropic has been explicit about the nuance: it hasn't determined that Claude Opus 4 definitively crosses the capability threshold that would require ASL-3. What it has concluded is that, for the first time, it can't rule that out with the same confidence it had for every previous model, given the continuous improvement in CBRN-related knowledge and capabilities. That's why the company calls this a "precautionary and provisional" measure: it's activating the stricter standard while it studies the actual risk level in more detail. Anthropic has ruled out the need for Opus 4 to operate under the higher ASL-4 level, and has also ruled out the need for Sonnet 4 to operate under ASL-3.

In practice, ASL-3 combines two fronts. On security, it hardens defenses against model weight theft by sophisticated non-state attackers. On deployment, it focuses — deliberately narrowly — on preventing Claude from assisting with complete, end-to-end workflows for developing CBRN weapons, without blocking general questions or information that's already public, such as the chemical formula for sarin. To that end, Anthropic has implemented "Constitutional Classifiers," real-time filters trained on synthetic data that monitor the model's inputs and outputs and block a specific category of harmful CBRN content.

Anthropic acknowledges that evaluating dangerous capabilities in an AI model is inherently difficult, and that as models approach concerning thresholds, it takes longer to determine their true status. That's why the company is choosing to activate the higher standard ahead of time: it lets Anthropic iterate on its defenses using real deployment experience rather than waiting for absolute certainty. If it later concludes that Opus 4 didn't actually cross the ASL-3 threshold, the company could withdraw or adjust these protections.

The underlying message is an uncomfortable one for the industry: the very company that sets the reference standard for AI safety is admitting it can no longer guarantee, with its usual confidence, that its most powerful models won't help build weapons of mass destruction. This is the first time that's happened — and it likely won't be the last.

Turning the headline into a check

The first step is to freeze the system's identity. Anthropic introduced Claude Opus 4 and Claude Sonnet 4 and activated ASL-3 measures for Opus. A commercial name may cover different revisions, automatic routes and tools. A test record should preserve date, access mode, configuration, permissions and the full output. Without that snapshot, an improvement or failure observed today cannot rigorously be attributed to the version another person will use tomorrow.

Next, turn how to separate measured capability, commercial claim, risk threshold and applied protection into cases with acceptance criteria. Build a local sample containing easy, ambiguous, long and deliberately impossible tasks. Record the input, the information the model may consult and what outcome would count as sufficient. A vendor-selected demonstration shows possibility; a test set preserved by the user measures reliability.

Autonomy needs a permission ladder. Reading and proposing are not the same as editing, sending or buying. A safer setup begins with read-only access, requires a preview and reserves execution for explicit approval. It also keeps a log and a rollback path. Judge the model by the errors the surrounding system contains, not by the confidence of its plan.

What the record must preserve

Cost and quality must be measured together. A cheap answer that must be reviewed from scratch can cost more than a slower but verifiable one. Measurement includes waiting, retries, consumption, human oversight and the consequences of failure. activation was precautionary and provisional; it did not confirm that the model had crossed the threshold. That boundary turns the announcement into a testable hypothesis rather than a promise to be believed.

An evidence sheet separates four columns: what the source claims, what it shows, what it did not measure and what would change the conclusion. That discipline prevents an absence from becoming a promise and a condition from vanishing in summary. It also lets the story be updated without rewriting history from a later outcome.

Include a negative case before deciding. Find a situation where the system, rule, transaction or study does not meet the need and record the signal that would require stopping. Selected successes show that something can happen; the negative case reveals the boundary and lowers the cost of discovering it after deployment.

The skill that outlasts the announcement

A valid comparison preserves denominator and axis. It does not pit a point figure against an average, future capacity against installed capacity or a forecast against an observation. When two sources use similar language, reconstruct what they counted and over what period. If those differ, publish them as different measures instead of inventing a ranking.

This article was produced with artificial intelligence under human editorial oversight.

Share this article

This website uses cookies to improve the browsing experience. Cookie policy.

↑↓ navigate ↵ open esc close