Claude Opus and Sonnet 4 Are Here: What Can Anthropic's New Generation of AI Do?
In May, Anthropic released a new generation of Claude language models, specifically Claude Opus and Sonnet 4, and both models immediately attracted enormous interest. According to benchmark tests, they are currently the best AI models for coding!
What Is Claude AI?
Claude AI is a direct competitor to major players such as OpenAI's GPT and Google's Gemini. According to available information, however, Claude 4 now beats the competition in several disciplines and represents a genuine leap forward.

What's New in Claude 4?
The new generation of Claude language models delivers its greatest improvements in programming and reasoning. Whether you choose the basic Sonnet or the powerful Opus, you will immediately notice several major improvements:
- Improved coding capabilities and deeper mathematical knowledge: Claude 4's greatest leap forward can be seen in programming and mathematics. In the SWE-bench and Terminal-bench tests, both models decisively outperform their biggest competitors, o3, GPT-4.1, and Gemini 2.5. This makes them the best tools for developers to date. In the AIME 2025 benchmark, which tests mathematical knowledge, there was a significant improvement over the previous Claude 3.7 series.
- Extended thinking feature: The new Claude 4 generation also introduces an “extended thinking mode,” a feature for extended thinking and advanced reasoning. Both models also use web search for this feature. The AI chat therefore not only provides thoughtful, multi-step answers, but also supports them with verified web data.
- Excellent memory: Claude 4 has a context window of 200k tokens. It can therefore analyze extensive documents. In addition, it has improved memory and remembers information across conversations. This makes it easier to work within the context of previous tasks.
- Better text comprehension and greater naturalness: Claude 4 handles the nuances of language much better than before. It understands context, responds more naturally, and its outputs require minimal human intervention—at least for now in the English version of the model.
Opus and Sonnet 4 are also much more reliable models than their predecessors. Thanks to extended thinking, they hallucinate less, are more self-critical, and use a final summary to reveal the logic behind their approach to the user—leading to more sophisticated conclusions and greater consistency.

Opus and Sonnet: What Is the Main Difference?
Both new models operate in basic and extended “reasoning” modes. While Opus 4 is the flagship and the “most powerful model” in the Claude 4 family, Sonnet 4 is an accessible and versatile tool for everyday use. Both, however, excel in programming, mathematics, and foreign language capabilities.
Claude 4: An Overall Improvement or a Sharp Lead in Only Certain Areas?
Although it might seem that Gemini is an utterly revolutionary AI with no competition, that is not the case. Certainly not everywhere! Although Opus 4 and Sonnet 4 achieved convincing results in programming and the MMMLU benchmark (language skills), they stagnated or saw a slight decline in many other areas.
Stagnation and Decline in MMMU and GPQA
In the MMMU test, Claude still lags behind its rivals and has even declined compared to its predecessor, Claude 3.7 Sonnet. It also shows similar results in the GPQA benchmark, which measures the level of complex analytical skills expected of a university student.
Worse Czech
Czech users are also encountering a decline in the model's Czech language proficiency. However, this can be attributed to the fact that it is a completely new model—its human-like and natural communication has improved considerably compared to its predecessors, which may temporarily affect the accuracy of its Czech translations.
Claude 4 vs. GPT or Gemini: Which Is More Creative and Which Is Better at Math?
The world of AI is incredibly dynamic, and every company is constantly advancing and improving its models. It is therefore impossible to determine which language models are the best—each excels at something different. However, if you are unsure whether OpenAI's GPT-4.1 or Gemini 2.5 Pro is better for your work, our table provides a brief overview of what each model excels at:

Is Claude 4 Worth Trying?
Definitely. Claude 4 will surprise you with how naturally it communicates and how masterfully it handles difficult and complex tasks (reasoning mode). If you enjoy trying new AI language models—or, conversely, are looking for your first one that communicates like a human and can handle tasks across different fields—Claude 4 is an excellent choice.



