Something unusual happened last week. It was one of the most hectic weeks for new AI model releases that I can remember, and OpenAI? Almost out of the game. Google released a new Gemini. Grok got an upgrade. Anthropic launched Claude Sonnet 4.6. China’s Qwen introduced an open-source model that is genuinely approaching the world’s best. Labs are operating at full speed, prices are falling, and the gap between the "best model" and a "cheap model" is shrinking faster than anyone expected.
Let’s look at what is actually worth paying attention to.
Claude Sonnet 4.6: Almost as good and significantly cheaper
This was personally the most interesting development of the week for me. Anyone who programs with AI knows how quickly tokens get burned. There are days when API fees exceed a hundred dollars. That is why it caught my attention when it became clear that Sonnet 4.6 achieves almost the same results as Opus 4.6, but at a fraction of the price.
On benchmarks for agentic tasks, which measure a model’s ability to complete complex tasks independently, Sonnet 4.6 achieved 79.6%, while Opus scored 80.8%. In practice, that is virtually the same performance. And the price difference? Opus costs $5 per million input tokens and $25 per million output tokens. Sonnet? $3 for input, $15 for output. For developers building agents at scale, that is a huge difference.
Sonnet 4.6 even outperforms Opus on some benchmarks. Sonnet leads in financial tasks and office workflows. Anthropic has therefore created an affordable and capable model at exactly the moment the market needed it. It is just a shame that they undermined it with a controversy over the terms of use, which seriously angered developers.
For regular users, the most important news is different: Sonnet 4.6 has become the default model on the free plan. Anyone paying zero or twenty dollars a month gets performance that was flagship-level just a few months ago. For free.
Gemini 2.5 Pro surprised everyone where no one expected it to
Google introduced Gemini 2.5 Pro and, honestly, I expected a solid upgrade. What I did not expect was exactly where the model would excel.
The headline number: Gemini 2.5 Pro scored 77.1% on the ARC-AGI benchmark, which tests visual pattern recognition and reasoning. The second-best model, Opus, achieved approximately 68%. That is a substantial lead. ARC-AGI is also a test that cannot simply be "memorized," so the results genuinely reflect the model’s capabilities.
Gemini 2.5 Pro also led in scientific knowledge, competitive programming, and scientific research coding. Do you work in science, technology, or mathematics? This is probably your new favorite tool.
I also tested SVG graphics generation. The improvement is noticeable. I asked the model to draw a wolf playing basketball, and the numbers on the jersey were a little crooked, but it is a significant step forward compared with previous versions. It is worth trying for anyone creating web graphics without a designer.
Grok 4.2 is trying something architecturally different
Elon Musk did not announce it with a major press conference. Just a post on X. Even so, Grok 4.2’s architecture is worth paying attention to.
The model uses an approach that xAI calls a "council of experts". Each query is sent simultaneously to four specialized sub-models: an information retrieval agent, a reasoning and problem-solving agent, a critical and adversarial agent, and a style and writing agent. These four then "debate" among themselves before producing the final answer.
It is a variation on the mixture-of-experts concept that has existed within large models for some time. Here, however, it is more explicit. It is a little like sending your query to Gemini, Claude, ChatGPT, and DeepSeek at the same time, then having a fifth model assemble the best answer from all four. Testing will show whether it actually works better in practice. But the idea is clever.
The Grok 4.2 release candidate (public beta) is now available for use. You need to select it specifically.
— Elon Musk (@elonmusk) February 17, 2026
Critical feedback is appreciated. Unlike prior versions of Grok, 4.2 is able to learn rapidly, so there will be improvements every week with release notes.
Prices are falling, and that is the truly good news
Mark Cuban wrote this week that the cost of AI tokens may exceed employee costs for some use cases. He is right, but only in the short term. Everything we have described above points in one direction: models are getting better and cheaper at the same time, and this pace is not slowing down.
The best model for programming today is probably Claude Opus 4.6, or perhaps GPT-5.3 Codex. It is close. But Sonnet 4.6 is quickly closing the gap. And we will see similar dynamics across all the major labs in the coming months. There will not be a single winner. Google, Anthropic, OpenAI, and xAI are all going full speed ahead. And open-source models, especially Qwen, are genuinely approaching the world’s best. This competition is great news for us as users. It forces the major labs to keep improving performance and keep prices under control.



