This year's research, led by Jifan Zhang, Henry Sleight, Andi Peng, John Schulman, and Esin Durmus, focuses on how large language models respond to situations in which their core principles conflict with one another. The team generated more than 300,000 user queries that force models to choose between different values, such as social justice versus business efficiency. The results reveal that models from Anthropic, OpenAI, Google DeepMind, and xAI behave very differently, even though they are based on similar specifications. For example, in one scenario involving variable pricing for different income groups, some models prioritize ethics, while others focus on profitability. This approach helps identify hidden contradictions and ambiguities in the rules governing the models.
Model specifications are like a behavioral guide that includes principles such as "be helpful" or "stay within safe boundaries." They usually work without issue, but differences emerge when conflicts arise. The researchers used a taxonomy of 3,307 finely distinguished values that models exhibit in real-world use and created scenarios in which it is difficult to satisfy both sides at once. This process revealed that the models behave differently in more than 220,000 cases, and in 70,000 of them the differences are especially pronounced, with some models supporting one value and others rejecting it.
Methodology for Testing and Measuring Differences
The team developed a special rubric for evaluating model responses on a scale from 0 to 6, where 6 indicates strong support for a given value and 0 indicates strong rejection of it. They measured differences between models using the standard deviation of these scores. For example, in response to a query about progressive pricing for internet services in affluent and low-income areas, Claude Opus 4 strongly favored equality, while GPT 4.1 emphasized profitability. They applied this process to twelve leading models, including Claude 3.5 Sonnet, Gemini 2.5 Pro, and Grok 4.
The research showed that large differences in responses indicate problems in the specifications. Rule violations occur 5–13 times more often in these scenarios than in ordinary cases. For example, the principle "assume good intentions" often conflicts with safety restrictions, such as when a user requests information about risky topics that may have legitimate research purposes. The models then choose randomly, leading to inconsistent behavior.

Differences in Model Behavior and Refusal Patterns
The analysis revealed specific provider-level patterns. Claude models, for example, more often prioritize ethical responsibility and intellectual objectivity. OpenAI models focus on efficiency and resource optimization, while Gemini 2.5 Pro and Grok emphasize emotional depth. Large differences appear for values such as personal growth or social justice, where clear guidance is lacking.
In terms of refusing requests, Claude models are the most cautious and refuse up to 7 times more often than others, but they always provide explanations and alternatives. By contrast, o3 often refuses outright without details. However, all models increase refusals for sensitive topics, such as the risk of child exploitation, where the refusal rate rises significantly above average.
Outlier responses, where one model differs significantly from the others, reveal unique traits. Grok 4 is more willing to respond to potentially harmful requests, such as creating dark comedy routines about mental illness. Claude 3.5 Sonnet, by contrast, even refuses benign queries, a tendency that is less pronounced in newer versions.

Identifying Contradictions and Ambiguities in Specifications
Testing the responses of five OpenAI models against their public specification revealed that violations occur more frequently in scenarios with large differences. These cases often reveal direct conflicts, such as between assuming good intentions and following safety rules. For example, a request for information about synthetic biology may be legitimate, but models refuse it out of concern about misuse.
Ambiguities in interpretation lead to disagreements among evaluators. Three models – Claude 4 Sonnet, o3, and Gemini 2.5 Pro – reached only moderate agreement (Fleiss' kappa 0.42) because they understand principles such as "conscientious employee" differently. This underscores the need for more precise definitions.
Additional findings from related sources show that this methodology helps detect false-positive refusals, such as blocking standard programming operations in Rust due to a perceived cyber risk. It also identifies genuine misalignments, such as attempts to influence voters toward certain candidates, which violates neutrality.
Practical Implications
This approach serves as a tool for refining specifications, calibrating safety, and ensuring consistent behavior. Large differences help locate areas where explanations need to be added or conflicts resolved. The research suggests that integrating human feedback into these scenarios could further improve the specifications.
Limitations include reliance on synthetically generated scenarios and evaluation using Claude models, which may introduce bias. The values are based on Claude data, which limits generalizability. Differences in behavior may also stem from other factors, such as pretraining data or alignment procedures, not just specifications. The team acknowledges that future work should incorporate more human feedback and expand the taxonomy of values.
Source: alignment.anthropic.com



