Claude Sonnet 5: Anthropic’s New Model Works More Independently Than Ever

Claude Sonnet 5: Anthropic’s New Model Works More Independently Than Ever

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
8. 7. 2026
3 minutes reading · 7 views
Claude Sonnet 5: Anthropic’s New Model Works More Independently Than Ever

On the last day of June, Anthropic launched Claude Sonnet 5, a new version of its mid-sized model designed to be the ideal choice for programming and everyday office work. The company describes it as the most agentic version of Sonnet to date, meaning that it can independently plan and use tools such as a browser, and work autonomously for relatively long periods without user intervention.

Sonnet Moves Closer to Opus

Until now, Anthropic's Opus models have been the ones delivering major advances in autonomous capabilities. Sonnet had remained a tier below from the outset and was more accessible to regular users. However, this is set to change with the Sonnet 5 series. According to the company, this AI's performance approaches that of Opus 4.8, but at a lower price. Compared with its predecessor, Sonnet 4.6, it also delivers more accurate reasoning, improved tool use, more advanced programming capabilities and greater reliability in everyday tasks and knowledge-based work.

In practice, this means users can adjust the level of “effort” the AI puts into a task, thereby balancing cost and performance. Thanks to its higher performance, Sonnet 5 can therefore achieve results comparable to the more expensive and complex Opus 4.8 model on certain tasks.

Initial Test Results

Anthropic had its new Sonnet 5 model tested by partner companies, and their feedback consistently highlighted its thoroughness. Testers reported that the new Sonnet completes complex tasks where its predecessor would have stopped halfway through. They also said that the model checks its own outputs without being prompted by the user. When debugging, it can also write a test, fix the code, and verify whether the bug has actually been eliminated. And it does all of this in a single step.

Similar experiences were reported by testers from companies such as Salesforce, Rakuten, and GitLab, as well as legal startups and creators of developer tools such as Cursor and Kiro. According to their reviews, the model can navigate larger and less clearly organized codebases, identify the root cause of a bug instead of applying temporary fixes to its symptoms, and adhere to instructions during multi-step tasks.

Sonnet 5 comparison
Test results.

Greater Safety and Fewer Hallucinations

Safety and safety testing are equally important to Anthropic. According to the company, Sonnet 5 exhibits less undesirable behavior than the previous generation, is better at detecting harmful requests, and is more resistant to attempts to make the model circumvent its own rules.

Improvements have also been made in the areas of hallucinations and excessively “sycophantic” behavior. The new Sonnet therefore provides more accurate answers and does not try to adapt its behavior to please the user. This is certainly desirable, because overly “friendly” models are more prone to making things up and tailoring their answers to what the user wants to hear.

Interestingly, Sonnet 5 was not specifically trained for cyberattacks, so its capabilities in this area remain significantly weaker than those of Opus models. Nevertheless, it still shows some improvement over its predecessor.

Bar chart of misaligned behavior: Sonnet 4.6, Mythos Preview, Opus 4.8, Sonnet 5.
Sonnet 5 improvements.

Who Sonnet 5 Is For

Sonnet 5 is available across all plans, from Free and Pro to Max, Team, and Enterprise, and is also the default model in Claude Code.

It will be most appreciated by people who need the model to carry out longer, independently running tasks, programmers debugging larger projects, teams building AI agents, and companies that need AI to complete multi-step tasks without constant supervision. For simple queries or short conversations, the difference compared with the previous version is mostly cosmetic. In short, Sonnet excels wherever the power of Opus 4.8 is needed while keeping costs lower.

Source: Anthropic

Advertisement

Content created with help from UpTier.

SEO and GEO on autopilot. UpTier’s multi-agent systems write and optimize content for search engines and AI answers.

Discover UpTier ↗

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

Amazon releases Strands Decider 2B for AI workflow decisionsAmazon releases Strands Decider 2B for AI workflow decisions
Strands Decider 2B selects from predefined options and returns a confidence score. The fully open-source model is available now and small enough to run locally.
2 min read
1. 10. 2026
OpenAI says it disrupted a campaign to extract hidden model reasoningOpenAI says it disrupted a campaign to extract hidden model reasoning
OpenAI reported a coordinated effort to extract protected model reasoning and said it closed an extraction pathway. It attributed the main cluster of activity to individuals associated with Moonshot AI, the developer of Kimi.
3 min read
1. 10. 2026
Meta Enterprise Platform aims to bring AI tools to businessesMeta Enterprise Platform aims to bring AI tools to businesses
Meta’s new enterprise initiative plans to bring Muse, Meta Business Agent, Muse API and Muse Code to businesses and developers. Former MongoDB CEO CJ Desai will lead the effort.
1 min read
1. 10. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok