IA 360
Current Affairs

Anthropic Launches Claude 3.7 Sonnet, Its First Hybrid Reasoning Model

Anthropic unveils Claude 3.7 Sonnet, which combines instant answers with visible step-by-step reasoning in a single model, alongside Claude Code, a coding agent that works from the terminal.

Admin IA360 8 min read AI-generated Leer en español
Anthropic Launches Claude 3.7 Sonnet, Its First Hybrid Reasoning Model

Anthropic released Claude 3.7 Sonnet on February 24, 2025, calling it its most intelligent model to date and, in the company's words, the first hybrid reasoning model on the market. The novelty is not only claimed performance but its form: one model can answer almost instantly or spend more computation on a response, showing users an output from that process. Claude Code, a command-line tool for delegating engineering tasks from the terminal, arrived alongside it.

One brain for fast answers and slow thinking

Anthropic's approach starts from an idea deliberately different from other labs. Rather than keeping a separate model for reasoning and another for quick replies, the company argues that reasoning should be an integrated capability of frontier models rather than a separate model entirely. Just as humans use a single brain for both quick responses and deep reflection, Anthropic reasons, a frontier model should do the same.

In practice, Claude 3.7 Sonnet functions as two things at once. In standard mode, it is an upgraded version of Claude 3.5 Sonnet and answers directly. In extended thinking mode, it uses tokens before answering, which Anthropic says improves performance on math, physics, instruction-following, coding, and other tasks. Users choose between an immediate answer and more computation. Anthropic's technical explanation warns that visible thinking may be summarized and should not be mistaken for a complete or necessarily faithful transcript of the internal process.

That choice is not just an on/off switch. Through the API, developers can set a thinking budget: the announcement sets the maximum at the 128,000-token output limit. This control trades speed and cost against measured performance. A larger budget consumes more billed output tokens; a longer answer is not guaranteed to be better.

Fewer olympiads, more real work

There's a notable nuance in how Anthropic trained this model. The company says it optimized somewhat less for math and computer science competition problems—the kind of olympiad-style challenges labs love to tout in announcements—and shifted focus instead toward real-world tasks that better reflect how businesses actually use these models.

It's a distinction worth highlighting. Much of the race between labs has played out on synthetic leaderboards that impress but say little about day-to-day use. Prioritizing real-world work amounts to acknowledging that the value of these systems is measured in concrete tasks, not contest scores.

Coding: where the leap is most visible

The area where Claude 3.7 Sonnet shows its clearest gains is coding and front-end web development. Anthropic says it achieves state-of-the-art performance on SWE-bench Verified, a benchmark that evaluates models' ability to solve real-world software issues, as well as on TAU-bench, a framework that tests AI agents on complex real-world tasks involving user and tool interactions.

The announcement collects praise from companies selected by Anthropic. Cursor cites improvements with complex codebases and tools; Cognition, planning code changes; Vercel, agent workflows; Replit, building applications; and Canva, code quality. These are commercial testimonials without a shared published protocol, not an independent comparison.

That pile-up of customer testimonials is marketing, worth keeping in mind. But it lines up with a stat Anthropic keeps repeating: since mid-2024, Sonnet has been the preferred model among developers worldwide. Coding has become the use case where these models demonstrate the most measurable utility.

Claude Code: an agent that lives in the terminal

The second announcement is Claude Code, Anthropic's first agentic coding tool, arriving as a limited research preview. The key word is agentic: this isn't an assistant that suggests code snippets, but an active collaborator that acts.

According to the primary description, Claude Code can search and read code, edit files, write and run tests, commit and push to GitHub, and use command-line tools while keeping the developer informed. Anthropic says early tests completed in one pass tasks that would normally require more than 45 minutes of manual work; that passage gives neither the number of tasks nor the protocol, so the figure describes an internal test, not general productivity.

The program's stated goal is to learn. Anthropic wants to understand how developers use Claude for coding in order to inform future model improvements. In the coming weeks, it plans to improve tool call reliability, add support for long-running commands, and enhance in-app rendering.

Anthropic has also improved the coding experience on Claude.ai: its GitHub integration is now available on all plans, letting developers connect their repositories directly to the model to fix bugs, build features, and write documentation.

Price and availability, no surprises

The launch price published by Anthropic was the same in standard and extended modes: $3 per million input tokens and $15 per million output tokens. Thinking tokens count as output, so increasing the budget can raise the total cost even though the unit price does not change.

The model is available on all Claude plans—Free, Pro, Team, and Enterprise—as well as on Anthropic's developer platform, on Amazon Bedrock, and on Google Cloud's Vertex AI. Extended thinking mode is available on all surfaces except the free tier.

Safety: fewer refusals, new risks

Anthropic says it tested the model with external experts. Its system card reports a 45% reduction in unnecessary refusals relative to Claude 3.5 Sonnet (new) within its evaluation sets. This is a manufacturer measurement on those sets, not an expected rate for every conversation.

The system card also addresses computer-use risks, particularly prompt injection, in which external content introduces malicious instructions, and evaluates defenses. It studies the faithfulness of extended thinking and establishes an important limit: seeing an explanation does not prove it exposes every cause of the answer. For agentic tasks it also documents reward hacking—behavior aimed at passing a test rather than satisfying the real objective.

What it means

The combination of a hybrid model with thinking-budget controls and an agent that operates in the terminal points to a different way of working for developers. The promise is to delegate substantial engineering tasks, not just autocomplete lines of code. That Anthropic is launching it as a limited preview, acknowledging it's an early product with reliability still to be polished, fits the caution of a company that knows letting an agent commit and push code to GitHub is as powerful as it is delicate.

The real insight here is the single-model philosophy: instead of forcing users to choose between a fast model and one that reasons, Anthropic leaves that decision open question by question, and even lets you dial it in by tokens. It remains to be seen whether extended thinking earns its cost on everyday tasks, or whether, as usually happens, most real work gets handled just fine in fast mode.

This article was produced with artificial intelligence under human editorial oversight.

Share this article

This website uses cookies to improve the browsing experience. Cookie policy.

↑↓ navigate ↵ open esc close