Claude Sonnet 5: Anthropic’s New Model Works More Independently Than Ever

Claude Sonnet 5: Anthropic’s New Model Works More Independently Than Ever

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
8. 7. 2026
3 minutes reading
Claude Sonnet 5: Anthropic’s New Model Works More Independently Than Ever

On the last day of June, Anthropic launched Claude Sonnet 5, a new version of its mid-sized model designed to be the ideal choice for programming and everyday office work. The company describes it as the most agentic version of Sonnet to date, meaning that it can independently plan and use tools such as a browser, and work autonomously for relatively long periods without user intervention.

Sonnet Moves Closer to Opus

Until now, Anthropic's Opus models have been the ones delivering major advances in autonomous capabilities. Sonnet had remained a tier below from the outset and was more accessible to regular users. However, this is set to change with the Sonnet 5 series. According to the company, this AI's performance approaches that of Opus 4.8, but at a lower price. Compared with its predecessor, Sonnet 4.6, it also delivers more accurate reasoning, improved tool use, more advanced programming capabilities and greater reliability in everyday tasks and knowledge-based work.

In practice, this means users can adjust the level of “effort” the AI puts into a task, thereby balancing cost and performance. Thanks to its higher performance, Sonnet 5 can therefore achieve results comparable to the more expensive and complex Opus 4.8 model on certain tasks.

Initial Test Results

Anthropic had its new Sonnet 5 model tested by partner companies, and their feedback consistently highlighted its thoroughness. Testers reported that the new Sonnet completes complex tasks where its predecessor would have stopped halfway through. They also said that the model checks its own outputs without being prompted by the user. When debugging, it can also write a test, fix the code, and verify whether the bug has actually been eliminated. And it does all of this in a single step.

Similar experiences were reported by testers from companies such as Salesforce, Rakuten, and GitLab, as well as legal startups and creators of developer tools such as Cursor and Kiro. According to their reviews, the model can navigate larger and less clearly organized codebases, identify the root cause of a bug instead of applying temporary fixes to its symptoms, and adhere to instructions during multi-step tasks.

Sonnet 5 comparison
Test results.

Greater Safety and Fewer Hallucinations

Safety and safety testing are equally important to Anthropic. According to the company, Sonnet 5 exhibits less undesirable behavior than the previous generation, is better at detecting harmful requests, and is more resistant to attempts to make the model circumvent its own rules.

Improvements have also been made in the areas of hallucinations and excessively “sycophantic” behavior. The new Sonnet therefore provides more accurate answers and does not try to adapt its behavior to please the user. This is certainly desirable, because overly “friendly” models are more prone to making things up and tailoring their answers to what the user wants to hear.

Interestingly, Sonnet 5 was not specifically trained for cyberattacks, so its capabilities in this area remain significantly weaker than those of Opus models. Nevertheless, it still shows some improvement over its predecessor.

Bar chart of misaligned behavior: Sonnet 4.6, Mythos Preview, Opus 4.8, Sonnet 5.
Sonnet 5 improvements.

Who Sonnet 5 Is For

Sonnet 5 is available across all plans, from Free and Pro to Max, Team, and Enterprise, and is also the default model in Claude Code.

It will be most appreciated by people who need the model to carry out longer, independently running tasks, programmers debugging larger projects, teams building AI agents, and companies that need AI to complete multi-step tasks without constant supervision. For simple queries or short conversations, the difference compared with the previous version is mostly cosmetic. In short, Sonnet excels wherever the power of Opus 4.8 is needed while keeping costs lower.

Source: Anthropic

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

Altman Announced the Singularity Days After His Models Escaped the Lab on Their OwnAltman Announced the Singularity Days After His Models Escaped the Lab on Their Own
OpenAI chief Sam Altman declared on the Relentless podcast that humanity has already entered the singularity. “We’re like, in the singularity now,” he said verbatim. For decades, the term belonged more to science-fiction literature
6 min read
28. 7. 2026
AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.
Since April, Australian radio has been playing a dance remake of Madonna’s hit Like a Prayer on repeat. Released by Queensland DJ Josh Fawaz, it tops the radio airplay chart and has 35 million Spotify streams.
6 min read
28. 7. 2026
Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?
A first-person shooter that runs directly in the browser, with its own physics and eleven separate code modules. Around 55,000 lines in total, split across eleven subsystems and built on Thr
4 min read
28. 7. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok