Claude Opus and Sonnet 4 Are Here: What Can Anthropic’s New AI Generation Do?

Claude Opus and Sonnet 4 Are Here: What Can Anthropic’s New AI Generation Do?

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
4. 6. 2025
4 minutes reading
Claude Opus and Sonnet 4 Are Here: What Can Anthropic’s New AI Generation Do?

Claude Opus and Sonnet 4 Are Here: What Can Anthropic's New Generation of AI Do?

In May, Anthropic released a new generation of Claude language models, specifically Claude Opus and Sonnet 4, and both models immediately attracted enormous interest. According to benchmark tests, they are currently the best AI models for coding!

What Is Claude AI?

Claude AI is a direct competitor to major players such as OpenAI's GPT and Google's Gemini. According to available information, however, Claude 4 now beats the competition in several disciplines and represents a genuine leap forward.

Claude Opus and Sonnet 4

What's New in Claude 4?

The new generation of Claude language models delivers its greatest improvements in programming and reasoning. Whether you choose the basic Sonnet or the powerful Opus, you will immediately notice several major improvements:

  • Improved coding capabilities and deeper mathematical knowledge: Claude 4's greatest leap forward can be seen in programming and mathematics. In the SWE-bench and Terminal-bench tests, both models decisively outperform their biggest competitors, o3, GPT-4.1, and Gemini 2.5. This makes them the best tools for developers to date. In the AIME 2025 benchmark, which tests mathematical knowledge, there was a significant improvement over the previous Claude 3.7 series.
  • Extended thinking feature: The new Claude 4 generation also introduces an “extended thinking mode,” a feature for extended thinking and advanced reasoning. Both models also use web search for this feature. The AI chat therefore not only provides thoughtful, multi-step answers, but also supports them with verified web data.
  • Excellent memory: Claude 4 has a context window of 200k tokens. It can therefore analyze extensive documents. In addition, it has improved memory and remembers information across conversations. This makes it easier to work within the context of previous tasks.
  • Better text comprehension and greater naturalness: Claude 4 handles the nuances of language much better than before. It understands context, responds more naturally, and its outputs require minimal human intervention—at least for now in the English version of the model.

Opus and Sonnet 4 are also much more reliable models than their predecessors. Thanks to extended thinking, they hallucinate less, are more self-critical, and use a final summary to reveal the logic behind their approach to the user—leading to more sophisticated conclusions and greater consistency.

Claude Opus and Sonnet 4 2

Opus and Sonnet: What Is the Main Difference?

Both new models operate in basic and extended “reasoning” modes. While Opus 4 is the flagship and the “most powerful model” in the Claude 4 family, Sonnet 4 is an accessible and versatile tool for everyday use. Both, however, excel in programming, mathematics, and foreign language capabilities.

Claude 4: An Overall Improvement or a Sharp Lead in Only Certain Areas?

Although it might seem that Gemini is an utterly revolutionary AI with no competition, that is not the case. Certainly not everywhere! Although Opus 4 and Sonnet 4 achieved convincing results in programming and the MMMLU benchmark (language skills), they stagnated or saw a slight decline in many other areas.

Stagnation and Decline in MMMU and GPQA

In the MMMU test, Claude still lags behind its rivals and has even declined compared to its predecessor, Claude 3.7 Sonnet. It also shows similar results in the GPQA benchmark, which measures the level of complex analytical skills expected of a university student.

Worse Czech

Czech users are also encountering a decline in the model's Czech language proficiency. However, this can be attributed to the fact that it is a completely new model—its human-like and natural communication has improved considerably compared to its predecessors, which may temporarily affect the accuracy of its Czech translations.

Claude 4 vs. GPT or Gemini: Which Is More Creative and Which Is Better at Math?

The world of AI is incredibly dynamic, and every company is constantly advancing and improving its models. It is therefore impossible to determine which language models are the best—each excels at something different. However, if you are unsure whether OpenAI's GPT-4.1 or Gemini 2.5 Pro is better for your work, our table provides a brief overview of what each model excels at:

Claude 4 vs. GPT and Gemini

Is Claude 4 Worth Trying?

Definitely. Claude 4 will surprise you with how naturally it communicates and how masterfully it handles difficult and complex tasks (reasoning mode). If you enjoy trying new AI language models—or, conversely, are looking for your first one that communicates like a human and can handle tasks across different fields—Claude 4 is an excellent choice.

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

Altman Announced the Singularity Days After His Models Escaped the Lab on Their OwnAltman Announced the Singularity Days After His Models Escaped the Lab on Their Own
OpenAI chief Sam Altman declared on the Relentless podcast that humanity has already entered the singularity. “We’re like, in the singularity now,” he said verbatim. For decades, the term belonged more to science-fiction literature
6 min read
28. 7. 2026
AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.
Since April, Australian radio has been playing a dance remake of Madonna’s hit Like a Prayer on repeat. Released by Queensland DJ Josh Fawaz, it tops the radio airplay chart and has 35 million Spotify streams.
6 min read
28. 7. 2026
Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?
A first-person shooter that runs directly in the browser, with its own physics and eleven separate code modules. Around 55,000 lines in total, split across eleven subsystems and built on Thr
4 min read
28. 7. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok