New Opus 4.5 Model Reigns Supreme in Coding

New Opus 4.5 Model Reigns Supreme in Coding

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
27. 11. 2025
3 minutes reading · 8 views
New Opus 4.5 Model Reigns Supreme in Coding

What Is Claude Opus 4.5?

Anthropic has just released its latest artificial intelligence (AI) model, called Claude Opus 4.5. This model was released on November 24, 2025, and is available on several platforms, including its apps, API, and major cloud services such as Amazon, Google, and Microsoft. People who have tested it say that Opus 4.5 can handle complex tasks without unnecessary hand-holding. For example, it can find and fix bugs in complex systems where earlier models failed. Feedback from customers such as the team at Replicate indicates that Opus 4.5 is now available at a price that makes it suitable for everyday tasks—it costs CZK 115 per million input tokens and CZK 575 per million output tokens.

The model is designed to excel at real-world tasks such as programming, working with Excel spreadsheets, and in-depth research. For example, in an internal Anthropic test in which engineering candidates are given a difficult take-home assignment, Opus 4.5 achieved a higher score than any human—all within a two-hour time limit. This means the model handled technical challenges faster and better, although it naturally lacks human skills such as collaboration or years of experience.

Benchmark Performance

Opus 4.5 excels in tests that measure real-world programming skills. It achieved 80.9% on the SWE-bench Verified benchmark, the highest score among all models. It also beat competitors such as GPT-5.1 at 77.9%, Sonnet 4.5 at 77.2%, Gemini 3 Pro at 76.2%, Opus 4.1 at 74.5%, and GPT-5.1 CodeMax at 73.3%. On multilingual SWE-bench, which tests code in eight languages, Opus 4.5 leads in seven of them.

Opus 4.5 benchmark results
Opus 4.5 benchmark results

In the long-running Vending-Bench test, the model operated longer without errors, achieving a 29% better result than Sonnet 4.5. Customers at Warp report that Opus 4.5 handles long autonomous tasks with 15% better results on their Terminal Bench.

In one example from the τ2-bench benchmark, the model creatively solved an airline ticket issue: instead of refusing a change to a basic economy ticket, it suggested upgrading the travel class first and then changing the flight, which is a legitimate solution under the rules.

SWE-bench results
SWE-bench results

Safety and Robustness

Anthropic focused on safety, and according to its system card, Opus 4.5 is the most strongly aligned model it has released. In tests for "concerning behavior," it achieved the lowest score among Anthropic's models, meaning fewer risks such as cooperating with human misuse or taking unintended actions. The model is more resistant to attacks such as prompt injection, in which hackers try to insert malicious instructions—Opus 4.5 has a 92% defense success rate, which is better than other models such as GPT-5.1 at 78% or Llama 3.1 at 65%.

This progress makes the model suitable for critical tasks where reliability is important, such as in enterprise environments. Details of all tests are available in the Claude Opus 4.5 system card.

Attack defense results
Attack defense results

Platform and Tool Updates

Opus 4.5 also brings updates to the Claude Developer Platform. The new "effort" parameter lets users control how many tokens the model uses—at the medium level, it achieves the same results as Sonnet 4.5 but with 76% fewer output tokens. At the high level, it outperforms Sonnet 4.5 by 4.3 percentage points while using 48% fewer tokens.

The model supports longer conversations through automatic context summarization and new tools such as zoom for computer use, which enables detailed screen viewing. Claude Code now includes Plan Mode, in which the model asks questions in advance and creates an editable plan in the plan.md file. The Claude app for Chrome enables work across browser tabs, and Claude for Excel is now available in beta to more users.

The related link also describes other models in the Claude 4.5 family: Sonnet 4.5 is best for complex agents and coding, with a context window of up to 1 million tokens in beta. Haiku 4.5 is the fastest, with performance close to Sonnet 4 at one-third of the price—CZK 23 per million input tokens and CZK 115 per million output tokens.

This model is available through the API under the designation claude-opus-4-5-20251101, and Anthropic plans further updates to keep pace with AI development. 

Advertisement

Content created with help from UpTier.

SEO and GEO on autopilot. UpTier’s multi-agent systems write and optimize content for search engines and AI answers.

Discover UpTier ↗

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

OpenAI gives Codex reusable cloud workspaces accessible from any deviceOpenAI gives Codex reusable cloud workspaces accessible from any device
Codex gains reusable cloud development environments, alongside voice controls in its CLI, code reviews in the ChatGPT desktop app and cloud-based security tools.
2 min read
2. 10. 2026
Amazon releases Strands Decider 2B for AI workflow decisionsAmazon releases Strands Decider 2B for AI workflow decisions
Strands Decider 2B selects from predefined options and returns a confidence score. The fully open-source model is available now and small enough to run locally.
2 min read
1. 10. 2026
OpenAI says it disrupted a campaign to extract hidden model reasoningOpenAI says it disrupted a campaign to extract hidden model reasoning
OpenAI reported a coordinated effort to extract protected model reasoning and said it closed an extraction pathway. It attributed the main cluster of activity to individuals associated with Moonshot AI, the developer of Kimi.
3 min read
1. 10. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok