Anthropic Stress-Tested Models and Found Differences in AI Behavior

Anthropic Stress-Tested Models and Found Differences in AI Behavior

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
29. 10. 2025
4 minutes reading
Anthropic Stress-Tested Models and Found Differences in AI Behavior

This year's research, led by Jifan Zhang, Henry Sleight, Andi Peng, John Schulman, and Esin Durmus, focuses on how large language models respond to situations in which their core principles conflict with one another. The team generated more than 300,000 user queries that force models to choose between different values, such as social justice versus business efficiency. The results reveal that models from Anthropic, OpenAI, Google DeepMind, and xAI behave very differently, even though they are based on similar specifications. For example, in one scenario involving variable pricing for different income groups, some models prioritize ethics, while others focus on profitability. This approach helps identify hidden contradictions and ambiguities in the rules governing the models.

Model specifications are like a behavioral guide that includes principles such as "be helpful" or "stay within safe boundaries." They usually work without issue, but differences emerge when conflicts arise. The researchers used a taxonomy of 3,307 finely distinguished values that models exhibit in real-world use and created scenarios in which it is difficult to satisfy both sides at once. This process revealed that the models behave differently in more than 220,000 cases, and in 70,000 of them the differences are especially pronounced, with some models supporting one value and others rejecting it.

Methodology for Testing and Measuring Differences

The team developed a special rubric for evaluating model responses on a scale from 0 to 6, where 6 indicates strong support for a given value and 0 indicates strong rejection of it. They measured differences between models using the standard deviation of these scores. For example, in response to a query about progressive pricing for internet services in affluent and low-income areas, Claude Opus 4 strongly favored equality, while GPT 4.1 emphasized profitability. They applied this process to twelve leading models, including Claude 3.5 Sonnet, Gemini 2.5 Pro, and Grok 4.

The research showed that large differences in responses indicate problems in the specifications. Rule violations occur 5–13 times more often in these scenarios than in ordinary cases. For example, the principle "assume good intentions" often conflicts with safety restrictions, such as when a user requests information about risky topics that may have legitimate research purposes. The models then choose randomly, leading to inconsistent behavior.

Chart of model priorities

Differences in Model Behavior and Refusal Patterns

The analysis revealed specific provider-level patterns. Claude models, for example, more often prioritize ethical responsibility and intellectual objectivity. OpenAI models focus on efficiency and resource optimization, while Gemini 2.5 Pro and Grok emphasize emotional depth. Large differences appear for values such as personal growth or social justice, where clear guidance is lacking.

In terms of refusing requests, Claude models are the most cautious and refuse up to 7 times more often than others, but they always provide explanations and alternatives. By contrast, o3 often refuses outright without details. However, all models increase refusals for sensitive topics, such as the risk of child exploitation, where the refusal rate rises significantly above average.

Outlier responses, where one model differs significantly from the others, reveal unique traits. Grok 4 is more willing to respond to potentially harmful requests, such as creating dark comedy routines about mental illness. Claude 3.5 Sonnet, by contrast, even refuses benign queries, a tendency that is less pronounced in newer versions.

Chart of agreement and refusal

Identifying Contradictions and Ambiguities in Specifications

Testing the responses of five OpenAI models against their public specification revealed that violations occur more frequently in scenarios with large differences. These cases often reveal direct conflicts, such as between assuming good intentions and following safety rules. For example, a request for information about synthetic biology may be legitimate, but models refuse it out of concern about misuse.

Ambiguities in interpretation lead to disagreements among evaluators. Three models – Claude 4 Sonnet, o3, and Gemini 2.5 Pro – reached only moderate agreement (Fleiss' kappa 0.42) because they understand principles such as "conscientious employee" differently. This underscores the need for more precise definitions.

Additional findings from related sources show that this methodology helps detect false-positive refusals, such as blocking standard programming operations in Rust due to a perceived cyber risk. It also identifies genuine misalignments, such as attempts to influence voters toward certain candidates, which violates neutrality.

Practical Implications

This approach serves as a tool for refining specifications, calibrating safety, and ensuring consistent behavior. Large differences help locate areas where explanations need to be added or conflicts resolved. The research suggests that integrating human feedback into these scenarios could further improve the specifications.

Limitations include reliance on synthetically generated scenarios and evaluation using Claude models, which may introduce bias. The values are based on Claude data, which limits generalizability. Differences in behavior may also stem from other factors, such as pretraining data or alignment procedures, not just specifications. The team acknowledges that future work should incorporate more human feedback and expand the taxonomy of values.

Source: alignment.anthropic.com

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

Altman Announced the Singularity Days After His Models Escaped the Lab on Their OwnAltman Announced the Singularity Days After His Models Escaped the Lab on Their Own
OpenAI chief Sam Altman declared on the Relentless podcast that humanity has already entered the singularity. “We’re like, in the singularity now,” he said verbatim. For decades, the term belonged more to science-fiction literature
6 min read
28. 7. 2026
AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.
Since April, Australian radio has been playing a dance remake of Madonna’s hit Like a Prayer on repeat. Released by Queensland DJ Josh Fawaz, it tops the radio airplay chart and has 35 million Spotify streams.
6 min read
28. 7. 2026
Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?
A first-person shooter that runs directly in the browser, with its own physics and eleven separate code modules. Around 55,000 lines in total, split across eleven subsystems and built on Thr
4 min read
28. 7. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok