On the last day of June, Anthropic launched Claude Sonnet 5, a new version of its mid-sized model designed to be the ideal choice for programming and everyday office work. The company describes it as the most agentic version of Sonnet to date, meaning that it can independently plan and use tools such as a browser, and work autonomously for relatively long periods without user intervention.
Sonnet Moves Closer to Opus
Until now, Anthropic's Opus models have been the ones delivering major advances in autonomous capabilities. Sonnet had remained a tier below from the outset and was more accessible to regular users. However, this is set to change with the Sonnet 5 series. According to the company, this AI's performance approaches that of Opus 4.8, but at a lower price. Compared with its predecessor, Sonnet 4.6, it also delivers more accurate reasoning, improved tool use, more advanced programming capabilities and greater reliability in everyday tasks and knowledge-based work.
In practice, this means users can adjust the level of “effort” the AI puts into a task, thereby balancing cost and performance. Thanks to its higher performance, Sonnet 5 can therefore achieve results comparable to the more expensive and complex Opus 4.8 model on certain tasks.
Initial Test Results
Anthropic had its new Sonnet 5 model tested by partner companies, and their feedback consistently highlighted its thoroughness. Testers reported that the new Sonnet completes complex tasks where its predecessor would have stopped halfway through. They also said that the model checks its own outputs without being prompted by the user. When debugging, it can also write a test, fix the code, and verify whether the bug has actually been eliminated. And it does all of this in a single step.
Similar experiences were reported by testers from companies such as Salesforce, Rakuten, and GitLab, as well as legal startups and creators of developer tools such as Cursor and Kiro. According to their reviews, the model can navigate larger and less clearly organized codebases, identify the root cause of a bug instead of applying temporary fixes to its symptoms, and adhere to instructions during multi-step tasks.
Greater Safety and Fewer Hallucinations
Safety and safety testing are equally important to Anthropic. According to the company, Sonnet 5 exhibits less undesirable behavior than the previous generation, is better at detecting harmful requests, and is more resistant to attempts to make the model circumvent its own rules.
Improvements have also been made in the areas of hallucinations and excessively “sycophantic” behavior. The new Sonnet therefore provides more accurate answers and does not try to adapt its behavior to please the user. This is certainly desirable, because overly “friendly” models are more prone to making things up and tailoring their answers to what the user wants to hear.
Interestingly, Sonnet 5 was not specifically trained for cyberattacks, so its capabilities in this area remain significantly weaker than those of Opus models. Nevertheless, it still shows some improvement over its predecessor.
Who Sonnet 5 Is For
Sonnet 5 is available across all plans, from Free and Pro to Max, Team, and Enterprise, and is also the default model in Claude Code.
It will be most appreciated by people who need the model to carry out longer, independently running tasks, programmers debugging larger projects, teams building AI agents, and companies that need AI to complete multi-step tasks without constant supervision. For simple queries or short conversations, the difference compared with the previous version is mostly cosmetic. In short, Sonnet excels wherever the power of Opus 4.8 is needed while keeping costs lower.
Source: Anthropic



