In recent months, we have witnessed a proliferation of reasoning AIs capable of breaking down tasks into steps and executing a thought process that culminates in more refined responses. The latest of these tools is Claude 3.7 Sonnet, the most advanced model developed by Anthropic. However, this AI has a particular feature: it is a hybrid reasoning model.
This means that, unlike other reasoning AIs, Claude 3.7 Sonnet allows the user to activate or deactivate these advanced thinking capabilities. Therefore, when you want the AI to provide a quick and simple response, you only need to request it, and the same applies when you need deep reflections for complex tasks. «We believe that reasoning should be an integrated capability of cutting-edge models rather than an entirely separate model», explains Anthropic.
In addition to this innovation, Anthropic has also introduced Claude Code, an agent coding tool currently in the preliminary phase. With Claude Code, developers can delegate complex engineering tasks to Claude directly from their terminal.
Introducing Claude 3.7 Sonnet: our most intelligent model to date. It’s a hybrid reasoning model, producing near-instant responses or extended, step-by-step thinking. One model, two ways to think. We’re also releasing an agentic coding tool: Claude Code. pic.twitter.com/jt7qQmFWuC — Anthropic (@AnthropicAI) February 24, 2025
Introducing Claude 3.7 Sonnet: our most intelligent model to date. It’s a hybrid reasoning model, producing near-instant responses or extended, step-by-step thinking.
One model, two ways to think.
We’re also releasing an agentic coding tool: Claude Code. pic.twitter.com/jt7qQmFWuC
— Anthropic (@AnthropicAI) February 24, 2025
Claude 3.7 Sonnet is the evolution of Claude 3.5 Sonnet (effectively, they have skipped 3.6) and is the first reasoning AI developed by Anthropic. Furthermore, it presents enhanced capabilities in mathematics, physics, instruction following, or coding, among other areas.
Being a hybrid model, it integrates the capabilities of a common LLM as well as those of a reasoning AI. The user simply needs to deploy the “Claude 3.7 Sonnet” button located in the AI’s text box and select the option “Normal” or “Extended”, depending on whether they require the tool to leverage its reasoning capabilities or not.
This feature distinguishes it from other reasoning models like OpenAI’s o3, which only allow complex, step-by-step thought processes to respond to inquiries. This implies that users have to choose a different model according to the task’s complexity. In fact, OpenAI recently announced its plans to eliminate the ChatGPT model selector and develop a tool capable of applying the most suitable model for each context, aiming to simplify and enhance the user experience.
This is also Anthropic’s objective. The AI laboratory, founded by former OpenAI employees, seeks to advance towards a model capable of deciding how long to “think” or “reason” a task, eliminating the intermediate step that obliges users to select the “Normal” or “Extended” mode.
Within its chat panel, Claude 3.7 Sonnet will display the internal reasoning process it undergoes to reach the final answer and will mark the time taken to arrive at its conclusion. However, Anthropic notes that it will not always reveal all its “thoughts” as some may be censored for security reasons.
The developer has optimized the thought modes of this artificial intelligence for real-world tasks that mirror how companies use this technology to boost productivity, such as coding problems or agency tasks.
Simultaneously, Claude 3.7 Sonnet has enhanced its ability to identify harmful requests and differentiate them from benign ones. Version 3.7 has reduced unnecessary rejection rates by 45% compared to its predecessor, 3.5.
Regarding the performance of this new model, in the SWE-Bench test (coding tasks) it revealed a precision of 62.3%, while OpenAI’s o3-mini AI achieved 49.3%. And, in the TAU-Bench test, which measures a model’s capability to interact with simulated users and external APIs, Claude 3.7 Sonnet achieved 81.2%, compared to OpenAI’s o1, which obtained 73.5%.
As an interesting detail, it is worth mentioning that Anthropic not only relied on these official tests to test its new AI, but also turned to other means such as playing the Pokemon Red video game on Game Boy.
For this purpose, as explained, «we equipped Claude with basic memory, pixel input on the screen, and function calls to press buttons and navigate the screen, which allowed it to play Pokemon continuously beyond its usual context limits, maintaining the gameplay over tens of thousands of interactions».
Claude 3.7 Sonnet managed to defeat three gym leaders and earn their badges. This was the best result achieved by a model in the Claude Sonnet family, whose first version Claude 3.0 Sonnet did not even manage to leave the house in Pallet Town.
«Pokemon is a fun way to appreciate Claude 3.7 Sonnet’s capabilities, but we hope these abilities will make a real-world impact far beyond gaming. The model’s ability to maintain focus and achieve open-ended objectives will aid developers in creating a wide range of next-generation AI agents», notes the developer.
Anthropic has granted access to Claude 3.7 Sonnet to all its users; however, not all plans provide access to its full version. The reasoning capability of the “Extended” mode will only be available to those who have subscribed to a paid plan (Pro, Team, or Enterprise). Meanwhile, the free plan will offer the enhanced capabilities of Claude 3.7 Sonnet as an LLM model but without the reasoning function.
It is also possible to use Claude 3.7 Sonnet and its reasoning capabilities in Anthropic API, Amazon Bedrock, and Vertex AI from Google Cloud.
When using this model via the API, «users can also control the thinking budget: they can instruct Claude to think for no more than N tokens, for any value of N up to its output limit of 128K tokens. This allows them to balance speed (and cost) with response quality».
The other innovation presented by Anthropic is Claude Code, its first active agent specialized in coding tasks. This model can «search and read code, edit files, write and run tests, commit and push code to GitHub, and use command-line tools, keeping the user informed at every step».
For now, Claude Code is in a limited research preview, but its capabilities have already shown great results. The developer explains that this model successfully completed tasks in a single pass that, typically, would take more than 45 minutes of manual work.
Anthropic has announced that it plans to implement continuous improvements based on user experience. Specifically, they plan to: «improve the reliability of tool calls, add support for long-running commands, enhance in-app representation, and expand Claude’s own understanding of its capabilities».
It is now possible to request access to the preliminary version of Claude Code by signing up for its waiting list.
Photo: Anthropic
Your email address will not be published. Required fields are marked *
Δ