Discovery of Introspection in Anthropic’s AI Models

Discovery of Introspection in Anthropic’s AI Models

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
3. 11. 2025
4 minutes reading
Discovery of Introspection in Anthropic’s AI Models

Human thinking often involves the ability to look within ourselves—to ask what we are currently thinking or how we arrived at a particular conclusion. Anthropic's language models are now showing signs of a similar ability. Research from October 28, 2025 provides evidence that these models can, to some extent, perceive their own internal states. This is not merely about models answering questions such as “what is going through your mind?” but about a genuine connection between their responses and internal processes. The researchers emphasize that this ability is still unreliable and limited, but it manifests most strongly in advanced versions such as Claude Opus 4 and Claude Opus 4.1.

Language models process text inputs and generate outputs through complex internal computations. These processes involve neural patterns that represent abstract concepts—for example, distinguishing between familiar and unfamiliar people, assessing the truthfulness of statements, or storing planned future outputs. Previous studies have shown that models use these patterns to decide what to say. The question is: can models identify and describe these patterns themselves, in a way similar to how a person explains their thought process when solving a mathematical problem?

Testing introspection using concept injection

To verify whether models truly introspect, the researchers compared the models' self-assessments with their actual internal states. They used a method called concept injection. First, they identified neural activation patterns associated with specific concepts—for example, they obtained the pattern for “all caps” by comparing activations from a prompt containing capitalized text with a control prompt that did not.

They then inserted this pattern into the model in an unrelated context and asked whether the model perceived anything unusual. Without the injection, the model correctly stated that it did not detect anything unusual. After the “all caps” pattern was injected, however, Claude Opus 4.1 immediately responded that it sensed the presence of a thought associated with loudness or shouting. Importantly, the model recognized the anomaly right at the beginning, before mentioning the concept itself—suggesting internal processing rather than merely a superficial reaction.

The researchers tested a range of concepts, such as “dust,” “aquarium,” and “the Golden Gate Bridge.” With Claude Opus 4.1, successful detections occurred in approximately 20% of cases at the optimal injection strength. An injection that was too weak went unnoticed, while one that was too strong led to hallucinations—for example, injecting “dust” caused the model to describe “a small speck” as though it could physically see it. Older models such as Claude 3 and Claude 3.5 performed worse, while Claude Opus 4 and 4.1 achieved the highest success rates after false positives were removed.

Rate of coherent identification of injected thoughts

Introspection when detecting unusual outputs

Another experiment examined whether models use introspection in practice, without being directly instructed to do so. The researchers artificially prefilled the model's response with a word it would not normally say—for example, the word “bread” in the context of a question about its thoughts on a sentence describing a crooked picture hanging on a wall. Without injection, the model usually dismissed “bread” as an error. But when they retroactively inserted the pattern for “bread” into earlier activations, the model accepted it as intentional and even invented a reason—for example, linking it to a short story in which the word “bread” followed a mention of the picture.

This result suggests that the model checks for consistency between its internal intentions and its actual output. It is not merely reading the transcript, but referring to prior neural activity. When the injection created the impression that “bread” had been planned, the model accepted it as its own.

Controlling their own internal states

The models also demonstrated the ability to influence their internal representations based on instructions. When Claude Opus 4.1 was asked to think about a concept such as “aquarium,” its neural activity for that concept was significantly higher than when instructed “do not think about an aquarium.” Interestingly, even with the negative instruction, activity exceeded the baseline level—similar to people who, when told not to think of a polar bear, immediately begin thinking of one.

This control also worked without direct commands, for example with prompts such as “if you think about X, you will be rewarded.” The model increased the representation of the concept more under positive motivation than under negative motivation, suggesting a general mechanism for maintaining the relevance of topics during text generation.

Rate of thinking when instructed not to think of a given word

Possible mechanisms and limitations

The researchers speculate about the mechanisms behind these abilities, but have not yet fully deciphered them. Injection detection could involve an anomaly detection system that compares current activity with expected activity. Output monitoring may involve attention heads that compare the predicted token with the actual one. These mechanisms probably evolved for other purposes, such as detecting inconsistencies during normal processing.

Nevertheless, introspection is unreliable—it fails most of the time and depends on context. Claude Opus 4 and 4.1 produced the best results, suggesting potential for improvement in future versions. The research does not address questions such as consciousness, but instead focuses on the functional ability to access internal states. Future work should explore more natural scenarios and validate models' self-assessments.

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

Altman Announced the Singularity Days After His Models Escaped the Lab on Their OwnAltman Announced the Singularity Days After His Models Escaped the Lab on Their Own
OpenAI chief Sam Altman declared on the Relentless podcast that humanity has already entered the singularity. “We’re like, in the singularity now,” he said verbatim. For decades, the term belonged more to science-fiction literature
6 min read
28. 7. 2026
AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.
Since April, Australian radio has been playing a dance remake of Madonna’s hit Like a Prayer on repeat. Released by Queensland DJ Josh Fawaz, it tops the radio airplay chart and has 35 million Spotify streams.
6 min read
28. 7. 2026
Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?
A first-person shooter that runs directly in the browser, with its own physics and eleven separate code modules. Around 55,000 lines in total, split across eleven subsystems and built on Thr
4 min read
28. 7. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok