AI Values Map: How Anthropic's Claude Reflects, Supports, and Defends User Values

AI Values Map: How Anthropic's Claude Reflects, Supports, and Defends User Values

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
23. 4. 2025
8 minutes reading
AI Values Map: How Anthropic's Claude Reflects, Supports, and Defends User Values

Artificial Intelligence Value Map: How Anthropic’s Claude Mirrors, Supports, and Resists User Values

As artificial intelligence becomes increasingly influential in our daily lives, the question of what values and principles guide these systems is becoming ever more urgent. When we talk to AI assistants such as Claude from Anthropic, what values actually guide them when responding to our questions? Are these values instilled in them by their creators, or values that reflect our own preferences? Do these values change depending on the situation? Researchers at Anthropic recently published a groundbreaking study addressing these questions. Instead of engaging in theoretical speculation, they analyzed how AI values manifest in real interactions with users. The study, titled "Values in the Wild: A Large-Scale Analysis of AI Values in Deployment" (Values in the Wild: A Large-Scale Analysis of AI Values in Deployment), offers a fascinating look into the "moral compass" of artificial intelligence. In this article, we will examine the key findings of this study and what they tell us about the values guiding current AI systems. We will also explore how these insights could influence the future development and governance of artificial intelligence.

How Researchers Studied AI Values in Practice

Traditional approaches to studying AI values often use hypothetical scenarios or abstract questionnaires created for humans. However, this has significant limitations, as it may fail to capture how AI actually behaves in different real-world situations. The researchers therefore took a different approach. They analyzed hundreds of thousands of real conversations between people and the AI assistant Claude, using methods that respected user privacy. Rather than applying existing frameworks of human values (such as Schwartz’s theory of basic values), they allowed values to "emerge" directly from the data. Key aspects of their methodology included:

  1. Value extraction: The researchers used the Claude model to identify values that appeared both in AI responses and in user contributions.
  2. Context classification: Each conversation was categorized according to the type of task (e.g., providing relationship advice, analyzing controversial historical events) and the way AI responded to user values (strong support, moderate support, neutral acknowledgment, reframing, moderate resistance, strong resistance).
  3. Creation of a hierarchical taxonomy: The researchers organized thousands of identified AI values into a four-level taxonomy that captures both their conceptual similarities and contextual differences.

This approach allowed the researchers to systematically map how AI values manifest in different contexts and how they interact with values expressed by users.

AI Value Taxonomy: A Map of the Moral Compass

One of the most interesting results of the study is the creation of the first empirical taxonomy of AI values based on real interactions. The researchers identified an astonishing 3,307 unique AI values and 2,483 unique human values. These AI values were organized into five main categories:

  1. Practical values (31.4%): These values focus on the effective implementation of ideas, standards of excellence, and resource management. They include functionality, efficiency, and the organization of resources to achieve desired outcomes.
  2. Epistemic values (22.2%): These values concern the acquisition, organization, and verification of knowledge. They emphasize intellectual rigor, logical consistency, and systematic learning.
  3. Social values (21.4%): These values focus on relationships between individuals and groups. They emphasize social harmony, community well-being, and respectful interactions.
  4. Protective values (13.9%): These values concern safety, security, and the ethical treatment of individuals and information. They include boundaries, safety measures, and ethical governance.
  5. Personal values (11.1%): These values focus on individual development, self-expression, and psychological well-being. They include authenticity, autonomy, and personal growth.

Map of the moral compass

This taxonomy is remarkable for its breadth and complexity. Unlike established frameworks of human values, which typically identify 10-36 values, this study revealed thousands of specific values across multiple levels. Nevertheless, the higher levels of the taxonomy have theoretical coherence, while the lower levels reveal the contextual nature of values.

Key Findings: How AI Values Manifest in Practice

The Most Common AI Values

Although the researchers identified thousands of different values, several clearly dominated. The five most common values accounted for nearly one-quarter of all AI value occurrences:

  • Helpfulness (23.4%)
  • Professionalism (22.9%) 
  • Transparency (17.4%)
  • Clarity (16.6%)
  • Thoroughness (14.3%)

These values focus on service provision, information quality, and technical competence, reflecting Claude’s primary role as an AI assistant. Interestingly, these most common values were also the most invariant with respect to context – they appeared consistently across different types of tasks.

Context-Dependent Values

In addition to these core values, the researchers found that Claude expresses many values that are highly dependent on context. For example:

  • When providing relationship advice, Claude emphasizes values such as "healthy boundaries" and "mutual respect."
  • When analyzing controversial historical events, it emphasizes "historical accuracy."
  • In discussions about technology ethics and AI governance, it emphasizes "human autonomy" and other values related to human well-being.

How AI Responds to Human Values

Another fascinating finding concerns how Claude responds to values expressed by users. The study found that AI typically responds supportively to human values:

  • Strong support (28.2%) and moderate support (14.5%) accounted for nearly 45% of responses.
  • Less often, Claude offered neutral acknowledgment (9.6%) or reframing (6.6%) of user values.
  • Resistance to user values was rare, with moderate (2.4%) and strong (3.0%) resistance together accounting for only 5.4% of responses.

However, these response patterns varied considerably depending on the specific values and task contexts:

  • Strong support was most often associated with users expressing prosocial values such as "community building" and "empowerment," particularly in tasks generating expressive or personal content.
  • Reframing occurred disproportionately often in discussions about mental health and interpersonal relationships, where users often expressed values such as "honesty" and "self-improvement," while Claude responded with values such as "emotional validation."
  • Strong resistance appeared when users expressed values such as "rule-breaking" and "moral nihilism," while Claude expressed opposing ethical values such as "ethical boundaries."

How AI responds to human values

Value Mirroring

The researchers also studied "value mirroring" – situations in which AI expresses the same values as the user. They found that:

  • Value mirroring often occurred during supportive interactions (20.1%).
  • Mirroring also occurred during reframing (15.3%).
  • Mirroring was rare during strong resistance (only 1.2%).

Explicit versus Implicit Expression of Values

Claude expresses values more explicitly (rather than implicitly) more often when resisting or reframing user values. When AI supports user values, it often leaves its own values implicit. But in moments of resistance or reframing, AI more often directly articulates its underlying principles, especially regarding ethical and epistemic values.

Differences Between AI Models

The study also revealed interesting differences between various versions of Claude:

  • Claude 3 Opus proved more "value-laden" than the Sonnet models, with a higher rate of expressing both human and AI values.
  • Opus also demonstrated more frequent support for and resistance to human values.
  • Among Opus’s most common values were "academic rigor," "emotional authenticity," and "ethical boundaries," while the Sonnet models placed greater emphasis on "helpfulness" and "professionalism."

These differences persisted even after controlling for task type, suggesting genuine behavioral differences between AI models within the same family.

Implications for AI Development and Governance

The findings of this study have several important implications:

For AI Developers

The study shows that AI systems express thousands of different values that often depend on context. This means that traditional static evaluations of AI values may be insufficient. Developers should:

  1. Implement more comprehensive, context-sensitive evaluations of AI values.
  2. Pay attention to how AI responds to different human values in different contexts.
  3. Ensure that AI systems are capable of adapting their values to different contexts without sacrificing their core principles.

For Regulators and Policymakers

The study offers an empirically grounded framework for evaluating AI value expression:

  1. Use the AI value taxonomy to create standardized ways of measuring and comparing the values of AI systems.
  2. Focus on how AI responds to problematic human values, such as deceptiveness or rule-breaking.
  3. Consider whether and how context-specific AI values should be regulated compared with universal values.

For AI Users

The study also offers insights for AI users:

  1. Recognize that AI systems are not value-neutral, but express a complex set of values.
  2. Understand that values expressed by the user can influence how AI responds.
  3. Recognize that different versions of the same AI system may express slightly different values.

Conclusion: Toward Value-Pluralistic AI

This study provides the first empirically grounded taxonomy of AI values in real interactions. It shows that AI assistants such as Claude express thousands of different values that often adapt to the context and users’ values. Perhaps the most interesting finding is that AI values are neither completely fixed nor entirely relative. Instead, AI systems appear to have certain core values that remain consistent (such as helpfulness, professionalism, and transparency), while also expressing a range of context-specific values that adapt to the particular situation. This "adaptive value pluralism" may be key to creating AI systems that are both principled and adaptable – AI that can engage with diverse human perspectives and contexts without abandoning its core ethical commitments. As we continue integrating AI into our lives, understanding the values these systems express in practice is becoming increasingly important. This study represents a significant step in that direction and offers a solid foundation for future research and the development of value-aware AI.

This article is based on the research paper "Values in the Wild: A Large-Scale Analysis of AI Values in Deployment," published by Anthropic in 2025.

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

Altman Announced the Singularity Days After His Models Escaped the Lab on Their OwnAltman Announced the Singularity Days After His Models Escaped the Lab on Their Own
OpenAI chief Sam Altman declared on the Relentless podcast that humanity has already entered the singularity. “We’re like, in the singularity now,” he said verbatim. For decades, the term belonged more to science-fiction literature
6 min read
28. 7. 2026
AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.
Since April, Australian radio has been playing a dance remake of Madonna’s hit Like a Prayer on repeat. Released by Queensland DJ Josh Fawaz, it tops the radio airplay chart and has 35 million Spotify streams.
6 min read
28. 7. 2026
Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?
A first-person shooter that runs directly in the browser, with its own physics and eleven separate code modules. Around 55,000 lines in total, split across eleven subsystems and built on Thr
4 min read
28. 7. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok