On February 5, 2026, Anthropic introduced Claude Opus 4.6, the latest version of its most powerful AI model. The new model brings significant improvements in programming, long-running agentic tasks, and working with large codebases. For the first time in the history of Opus-class models, it offers a 1-million-token context window in beta.
Improved programming capabilities
Claude Opus 4.6 excels primarily at coding. The model plans more carefully, can sustain agentic tasks for longer, and works more reliably across larger codebases. It now has better code review and debugging capabilities, allowing it to catch its own mistakes. The model achieved the highest score on Terminal-Bench 2.0, which tests agentic coding.
According to GitHub, the model handles complex, multi-step programming tasks that developers tackle every day, particularly agentic workflows requiring planning and tool calls. Replit confirmed that Opus 4.6 represents a massive leap in agentic planning—it breaks down complex tasks into independent subtasks, runs tools and subagents in parallel, and identifies obstacles with high accuracy.
Dominates performance benchmarks
Opus 4.6 ranks first in the Artificial Analysis Intelligence Index, a synthetic metric comprising 10 evaluations covering capabilities including agentic tasks, programming, and scientific reasoning. The model leads in three evaluations: GDPval-AA (real-world agentic tasks), TerminalBench (agentic coding and terminal use), and CritPT (research-level physics problems).
In the GDPval-AA test, which measures model performance on knowledge-work tasks ranging from preparing presentations and analyzing data to editing video, Opus 4.6 outperformed OpenAI's GPT-5.2 by approximately 144 Elo points and its predecessor, Claude Opus 4.5, by 190 points. The model also achieved the best performance on BrowseComp, which measures a model's ability to find hard-to-access information online.
On Humanity's Last Exam, a challenging multidisciplinary reasoning test, Opus 4.6 leads all other frontier models. In scientific reasoning, the model achieved the highest score of 13% on CritPT, an evaluation consisting of unpublished research-level physics problems.
Working with long context
Opus 4.6 has significantly improved its ability to retrieve relevant information from large document sets. On the MRCR v2 benchmark (8-needle 1M variant), which tests a model's ability to find information "hidden" within a vast amount of text, Opus 4.6 achieved a score of 76%, while Sonnet 4.5 achieved only 18.5%. This represents a qualitative leap in how much context the model can actually use while maintaining top-tier performance.
A common complaint about AI models is "context rot," where performance deteriorates once a conversation exceeds a certain number of tokens. In this respect, Opus 4.6 performs significantly better than its predecessors.
New features and capabilities
Anthropic introduced several key changes alongside the release of Opus 4.6. The model introduces a new "adaptive thinking" mode, which replaces the previous "extended thinking" mode. Instead of setting a thinking-token budget, developers can now control the model's thinking using the "effort" setting with four levels: low, medium, high, and maximum.
The maximum output token limit has doubled from 64,000 to 128,000 tokens compared with Opus 4.5. The model also supports a new context compaction feature in beta, which automatically summarizes and replaces older context when a conversation approaches a configurable threshold.
Pricing and availability
Claude Opus 4.6 is available on claude.ai, through Anthropic's API, and on all major cloud platforms, including Google Cloud Vertex, AWS Bedrock, and Microsoft Azure. Pricing remains the same as for Opus 4.5: CZK 125 per million input tokens and CZK 625 per million output tokens.
According to an analysis by Artificial Analysis, running the Intelligence Index with Opus 4.6 in adaptive thinking mode at maximum effort cost CZK 62,150 ($2,486), more than OpenAI's GPT-5.2 (xhigh), which cost CZK 57,600 ($2,304). However, the model used significantly fewer tokens than the competition—58 million output tokens, approximately twice as many as Opus 4.5, but substantially fewer than GPT-5.2 with xhigh reasoning effort (130 million).
Safety and reliability
According to Anthropic's extensive system card, Opus 4.6 has an overall safety profile as good as or better than any other frontier model in the industry, with low rates of misbehavior across safety evaluations. The model also has the lowest over-refusal rate of all recent Claude models, referring to cases in which the model fails to answer harmless queries.
Notion stated that Opus 4.6 is the strongest model Anthropic has released. It takes complex requests and actually completes them, breaking them down into concrete steps, carrying them out, and producing polished work even when the task is ambitious. Asana confirmed that Opus 4.6 is the best model it has tested to date, with exceptional reasoning and planning capabilities.



