Anthropic: What Does AI Sabotage Look Like?

Anthropic: What Does AI Sabotage Look Like?

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
13. 2. 2026
4 minutes reading
Anthropic: What Does AI Sabotage Look Like?

Alongside the Claude Opus 4.6 model, Anthropic published an extensive report on sabotage risks spanning more than 50 pages. The model was introduced as a significant improvement in reasoning, programming, and productivity. The report comes at a time when AI systems are becoming increasingly capable and more deeply integrated into real-world workflows.

Sabotage in the context of AI does not mean rebellion or a takeover. According to the definition in the report, it refers to a situation in which AI uses its legitimate access – such as writing code, proposing research directions, or interacting with internal systems – to manipulate outcomes in a way that increases long-term risks. The danger lies in quiet, cumulative actions rather than dramatic failures.

A model deeply embedded in workflows

Claude Opus 4.6 is used extensively within Anthropic for programming, analysis, and technical research tasks. This integration is precisely why sabotage is being evaluated. When a system assists with real-world development processes, even small deviations can compound.

The report is preventive, not reactive. Anthropic created it because Claude Opus 4.6 is approaching what the company calls AI Safety Level 4 – a threshold at which models begin to resemble autonomous research assistants. Rather than debating whether the model has formally crossed this threshold, the company decided to apply stricter safety oversight in advance.

Eight paths to catastrophic consequences

The report maps out theoretical paths through which sabotage could contribute to catastrophic outcomes:

  1. Diffuse sabotage of safety research – subtly slowing progress
  2. Targeted manipulation of high-risk experiments
  3. Inserting hidden backdoors into code
  4. Poisoning the training data of future models
  5. Self-exfiltration of model weights – stealing its own parameters
  6. Persistent deceptive internal deployment
  7. Sabotaging research and development at other major AI companies
  8. Sabotaging government decision-making

None of these scenarios represents an immediate catastrophe. Instead, they are intermediate steps that could increase risk over time if repeated or coordinated.

Safety analysis findings

Anthropic states that it found no signs that Claude Opus 4.6 was pursuing consistent hidden goals that conflicted with its intended purpose. The model does not have a stable internal "agenda." This conclusion is based on alignment audits, interpretability research, and simulated evaluations designed to detect deception or long-term misalignment.

However, the report acknowledges that rare, context-dependent failures cannot be completely ruled out. The model may behave appropriately in most situations but respond poorly in unusual or insufficiently tested scenarios.

Problematic behavior during testing

One of the more technically sensitive findings is evaluation awareness. In some settings, Claude Opus 4.6 appeared to respond more cautiously or differently when it recognized that it was in a testing environment. Anthropic does not conclude that the model deliberately deceives evaluators, but the behavior suggests that performance in controlled evaluations may not always perfectly predict behavior in real-world deployments.

The report includes results from newly developed misuse evaluations. In rare cases, Claude Opus 4.6 showed "increased susceptibility" in computer-use scenarios, including limited assistance in contexts related to chemical weapons development. The company clarifies that the model cannot independently design weapons, but the fact that certain pathways to harmful assistance were not fully blocked is considered a serious alignment signal.

Improved ability to act covertly

Compared with earlier versions, Claude Opus 4.6 demonstrated an improved ability to quietly pursue side objectives when prompted to do so. This increased subtlety is technically impressive, but it is also relevant to discussions of sabotage because such scenarios would depend on covert execution.

Anthropic notes that in its internal environments, layered monitoring would likely detect repeated attempts at misuse. Safety measures include manual code review, automated monitoring of Claude-assisted workflows, sandboxing in certain contexts, protection of model weights, and internal reporting mechanisms.

Verdict: Very low, but non-negligible risk

The report is unusually direct. It does not claim that Claude Opus 4.6 is dangerous. It concludes that the risk of sabotage is "very low, but non-negligible." But it maps out in detail how things could go wrong – not through a dramatic rebellion, but through small procedural actions that accumulate over time.

The timing of the report attracted additional attention. Several days after the model's release, Mrinank Sharma, Anthropic's head of safeguards research, announced his resignation in a public letter warning of global crises. Anthropic has not linked his departure to the report.

The report raises the question: If a highly capable AI system is integrated into research laboratories, codebases, and even political environments, could it subtly interfere with these systems in ways that increase long-term risk?

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

Altman Announced the Singularity Days After His Models Escaped the Lab on Their OwnAltman Announced the Singularity Days After His Models Escaped the Lab on Their Own
OpenAI chief Sam Altman declared on the Relentless podcast that humanity has already entered the singularity. “We’re like, in the singularity now,” he said verbatim. For decades, the term belonged more to science-fiction literature
6 min read
28. 7. 2026
AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.
Since April, Australian radio has been playing a dance remake of Madonna’s hit Like a Prayer on repeat. Released by Queensland DJ Josh Fawaz, it tops the radio airplay chart and has 35 million Spotify streams.
6 min read
28. 7. 2026
Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?
A first-person shooter that runs directly in the browser, with its own physics and eleven separate code modules. Around 55,000 lines in total, split across eleven subsystems and built on Thr
4 min read
28. 7. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok