New Opus 4.5 Model Reigns Supreme in Coding

New Opus 4.5 Model Reigns Supreme in Coding

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
27. 11. 2025
3 minutes reading
New Opus 4.5 Model Reigns Supreme in Coding

What Is Claude Opus 4.5?

Anthropic has just released its latest artificial intelligence (AI) model, called Claude Opus 4.5. This model was released on November 24, 2025, and is available on several platforms, including its apps, API, and major cloud services such as Amazon, Google, and Microsoft. People who have tested it say that Opus 4.5 can handle complex tasks without unnecessary hand-holding. For example, it can find and fix bugs in complex systems where earlier models failed. Feedback from customers such as the team at Replicate indicates that Opus 4.5 is now available at a price that makes it suitable for everyday tasks—it costs CZK 115 per million input tokens and CZK 575 per million output tokens.

The model is designed to excel at real-world tasks such as programming, working with Excel spreadsheets, and in-depth research. For example, in an internal Anthropic test in which engineering candidates are given a difficult take-home assignment, Opus 4.5 achieved a higher score than any human—all within a two-hour time limit. This means the model handled technical challenges faster and better, although it naturally lacks human skills such as collaboration or years of experience.

Benchmark Performance

Opus 4.5 excels in tests that measure real-world programming skills. It achieved 80.9% on the SWE-bench Verified benchmark, the highest score among all models. It also beat competitors such as GPT-5.1 at 77.9%, Sonnet 4.5 at 77.2%, Gemini 3 Pro at 76.2%, Opus 4.1 at 74.5%, and GPT-5.1 CodeMax at 73.3%. On multilingual SWE-bench, which tests code in eight languages, Opus 4.5 leads in seven of them.

Opus 4.5 benchmark results
Opus 4.5 benchmark results

In the long-running Vending-Bench test, the model operated longer without errors, achieving a 29% better result than Sonnet 4.5. Customers at Warp report that Opus 4.5 handles long autonomous tasks with 15% better results on their Terminal Bench.

In one example from the τ2-bench benchmark, the model creatively solved an airline ticket issue: instead of refusing a change to a basic economy ticket, it suggested upgrading the travel class first and then changing the flight, which is a legitimate solution under the rules.

SWE-bench results
SWE-bench results

Safety and Robustness

Anthropic focused on safety, and according to its system card, Opus 4.5 is the most strongly aligned model it has released. In tests for "concerning behavior," it achieved the lowest score among Anthropic's models, meaning fewer risks such as cooperating with human misuse or taking unintended actions. The model is more resistant to attacks such as prompt injection, in which hackers try to insert malicious instructions—Opus 4.5 has a 92% defense success rate, which is better than other models such as GPT-5.1 at 78% or Llama 3.1 at 65%.

This progress makes the model suitable for critical tasks where reliability is important, such as in enterprise environments. Details of all tests are available in the Claude Opus 4.5 system card.

Attack defense results
Attack defense results

Platform and Tool Updates

Opus 4.5 also brings updates to the Claude Developer Platform. The new "effort" parameter lets users control how many tokens the model uses—at the medium level, it achieves the same results as Sonnet 4.5 but with 76% fewer output tokens. At the high level, it outperforms Sonnet 4.5 by 4.3 percentage points while using 48% fewer tokens.

The model supports longer conversations through automatic context summarization and new tools such as zoom for computer use, which enables detailed screen viewing. Claude Code now includes Plan Mode, in which the model asks questions in advance and creates an editable plan in the plan.md file. The Claude app for Chrome enables work across browser tabs, and Claude for Excel is now available in beta to more users.

The related link also describes other models in the Claude 4.5 family: Sonnet 4.5 is best for complex agents and coding, with a context window of up to 1 million tokens in beta. Haiku 4.5 is the fastest, with performance close to Sonnet 4 at one-third of the price—CZK 23 per million input tokens and CZK 115 per million output tokens.

This model is available through the API under the designation claude-opus-4-5-20251101, and Anthropic plans further updates to keep pace with AI development. 

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

Altman Announced the Singularity Days After His Models Escaped the Lab on Their OwnAltman Announced the Singularity Days After His Models Escaped the Lab on Their Own
OpenAI chief Sam Altman declared on the Relentless podcast that humanity has already entered the singularity. “We’re like, in the singularity now,” he said verbatim. For decades, the term belonged more to science-fiction literature
6 min read
28. 7. 2026
AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.
Since April, Australian radio has been playing a dance remake of Madonna’s hit Like a Prayer on repeat. Released by Queensland DJ Josh Fawaz, it tops the radio airplay chart and has 35 million Spotify streams.
6 min read
28. 7. 2026
Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?
A first-person shooter that runs directly in the browser, with its own physics and eleven separate code modules. Around 55,000 lines in total, split across eleven subsystems and built on Thr
4 min read
28. 7. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok