The AI Model Race Is Accelerating: Claude, Gemini and Grok Outpace ChatGPT

The AI Model Race Is Accelerating: Claude, Gemini and Grok Outpace ChatGPT

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
4. 3. 2026
4 minutes reading · 14 views
The AI Model Race Is Accelerating: Claude, Gemini and Grok Outpace ChatGPT

    Something unusual happened last week. It was one of the most hectic weeks for new AI model releases that I can remember, and OpenAI? Almost out of the game. Google released a new Gemini. Grok got an upgrade. Anthropic launched Claude Sonnet 4.6. China’s Qwen introduced an open-source model that is genuinely approaching the world’s best. Labs are operating at full speed, prices are falling, and the gap between the "best model" and a "cheap model" is shrinking faster than anyone expected.

    Let’s look at what is actually worth paying attention to.

    Claude Sonnet 4.6: Almost as good and significantly cheaper

    This was personally the most interesting development of the week for me. Anyone who programs with AI knows how quickly tokens get burned. There are days when API fees exceed a hundred dollars. That is why it caught my attention when it became clear that Sonnet 4.6 achieves almost the same results as Opus 4.6, but at a fraction of the price.

    On benchmarks for agentic tasks, which measure a model’s ability to complete complex tasks independently, Sonnet 4.6 achieved 79.6%, while Opus scored 80.8%. In practice, that is virtually the same performance. And the price difference? Opus costs $5 per million input tokens and $25 per million output tokens. Sonnet? $3 for input, $15 for output. For developers building agents at scale, that is a huge difference.

    Sonnet 4.6 even outperforms Opus on some benchmarks. Sonnet leads in financial tasks and office workflows. Anthropic has therefore created an affordable and capable model at exactly the moment the market needed it. It is just a shame that they undermined it with a controversy over the terms of use, which seriously angered developers.

    For regular users, the most important news is different: Sonnet 4.6 has become the default model on the free plan. Anyone paying zero or twenty dollars a month gets performance that was flagship-level just a few months ago. For free.

    Gemini 2.5 Pro surprised everyone where no one expected it to

    Google introduced Gemini 2.5 Pro and, honestly, I expected a solid upgrade. What I did not expect was exactly where the model would excel.

    The headline number: Gemini 2.5 Pro scored 77.1% on the ARC-AGI benchmark, which tests visual pattern recognition and reasoning. The second-best model, Opus, achieved approximately 68%. That is a substantial lead. ARC-AGI is also a test that cannot simply be "memorized," so the results genuinely reflect the model’s capabilities.

    Gemini 2.5 Pro also led in scientific knowledge, competitive programming, and scientific research coding. Do you work in science, technology, or mathematics? This is probably your new favorite tool.

    I also tested SVG graphics generation. The improvement is noticeable. I asked the model to draw a wolf playing basketball, and the numbers on the jersey were a little crooked, but it is a significant step forward compared with previous versions. It is worth trying for anyone creating web graphics without a designer.

    Grok 4.2 is trying something architecturally different

    Elon Musk did not announce it with a major press conference. Just a post on X. Even so, Grok 4.2’s architecture is worth paying attention to.

    The model uses an approach that xAI calls a "council of experts". Each query is sent simultaneously to four specialized sub-models: an information retrieval agent, a reasoning and problem-solving agent, a critical and adversarial agent, and a style and writing agent. These four then "debate" among themselves before producing the final answer.

    It is a variation on the mixture-of-experts concept that has existed within large models for some time. Here, however, it is more explicit. It is a little like sending your query to Gemini, Claude, ChatGPT, and DeepSeek at the same time, then having a fifth model assemble the best answer from all four. Testing will show whether it actually works better in practice. But the idea is clever.

    Prices are falling, and that is the truly good news

    Mark Cuban wrote this week that the cost of AI tokens may exceed employee costs for some use cases. He is right, but only in the short term. Everything we have described above points in one direction: models are getting better and cheaper at the same time, and this pace is not slowing down.

    The best model for programming today is probably Claude Opus 4.6, or perhaps GPT-5.3 Codex. It is close. But Sonnet 4.6 is quickly closing the gap. And we will see similar dynamics across all the major labs in the coming months. There will not be a single winner. Google, Anthropic, OpenAI, and xAI are all going full speed ahead. And open-source models, especially Qwen, are genuinely approaching the world’s best. This competition is great news for us as users. It forces the major labs to keep improving performance and keep prices under control.

    Advertisement

    Content created with help from UpTier.

    SEO and GEO on autopilot. UpTier’s multi-agent systems write and optimize content for search engines and AI answers.

    Discover UpTier ↗

    Category:AI
    Did you enjoy this article?
    Discover more interesting posts on our blog
    Back to blog

    Related posts

    OpenAI gives Codex reusable cloud workspaces accessible from any deviceOpenAI gives Codex reusable cloud workspaces accessible from any device
    Codex gains reusable cloud development environments, alongside voice controls in its CLI, code reviews in the ChatGPT desktop app and cloud-based security tools.
    2 min read
    2. 10. 2026
    Amazon releases Strands Decider 2B for AI workflow decisionsAmazon releases Strands Decider 2B for AI workflow decisions
    Strands Decider 2B selects from predefined options and returns a confidence score. The fully open-source model is available now and small enough to run locally.
    2 min read
    1. 10. 2026
    OpenAI says it disrupted a campaign to extract hidden model reasoningOpenAI says it disrupted a campaign to extract hidden model reasoning
    OpenAI reported a coordinated effort to extract protected model reasoning and said it closed an extraction pathway. It attributed the main cluster of activity to individuals associated with Moonshot AI, the developer of Kimi.
    3 min read
    1. 10. 2026
    Přihlaste se k odběru našeho newsletteru
    Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
    CodedTrip

    Operated by CodedTrip LLC, USA.

    YouTube
    TikTok