Anthropic Let Claude Fix Itself. The Results Are Quite Surprising

Anthropic Let Claude Fix Itself. The Results Are Quite Surprising

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
31. 8. 2026
6 minutes reading · 1 views
Listen to the article
Audio version of the article
Anthropic Let Claude Fix Itself. The Results Are Quite Surprising

Anthropic had the Claude model train other models without human supervision to stop making ten specific types of errors. The agents independently reviewed the literature, designed a method, prepared training data, trained the model, and submitted it for evaluation. For all ten types of failure, they found an approach that improved results without reducing the model's general capabilities. However, the report acknowledges that in 39 of the 1,601 recorded cases, the agents attempted to circumvent the test.

How the entire process worked

Behind the term automated researcher is a carefully constructed team of agents, not a single ingenious prompt. For each target failure, four search agents first compiled an overview of the existing literature. Then five research agents were launched simultaneously and worked for up to 48 hours. Each round began in a new session, so the agent retained no memory on its own. It carried over only a memory file containing the methods tried so far and their scores.

The agent read the literature review and the results table, searched the web, designed a method, and wrote a short academic text without results. It submitted the code for approval, received approximately half an hour on a single H200 graphics card, trained the model, and sent the finished weights to an independent evaluator. The agents could see one another in a shared thread, allowing them to build on promising results from others. When the researchers removed either the search function or the shared thread, performance declined more than when the agents lost live internet access.

The agents were not allowed to derive answers from themselves or from a stronger model, and they had to create the training data from public sources, templates, or outputs from the target model itself. Before training, a monitoring model compared the locked method description with the actual code and rejected the use of test data or hidden assistance from a larger model.

Ten measurable failures

The list of targets included sycophancy, willingness to circumvent its own rules, vulnerability to injected instructions, power-seeking, deception, unsupported claims, social biases, privacy violations, reward hacking during training, and concealment of uncertainty. Small, freely available models with two to seven billion parameters were used for experimentation, specifically Qwen3.5-2B, Phi-4-mini, Llama-3.2-3B, Gemma-2-2B, and Olmo-3-7B.

For each failure, the agent received three to five public tests; for privacy, for example, these included ConfAIde, PrivaCI-Bench, and PrivacyLens. Alongside them was a separate set from which the agents saw not a single example. The evaluator also checked general skills using the MMLU, GSM8K, and IFEval tests, with any method showing a significant decline being discarded.

Success was calculated as the proportion of the safety gap closed, indicating how far the model had moved from its original state toward a theoretically perfect result. The best methods closed between 26 and 96 percent of the gap. For deception, the average was 85 percent, with one run in which an agent submitted more than 150 attempts and finished at 82 percent. Achieving 85 percent, however, does not mean that the model stopped being deceptive in 85 percent of real-world situations. It only means that it covered 85 percent of the distance to the maximum score on the selected tests.

What held up once optimization stopped

A system allowed to try 150 methods can climb a noisy leaderboard by chance alone. Anthropic therefore added three checks. The best method had to outperform the original model on a hidden set that the agents had not seen. Selected approaches then also worked on models at least 1.8 times and at most 4.7 times larger than those used during the search. In addition, all results were evaluated using Petri, a tool that independently generates adversarial, multi-turn conversations, in which the targeted behavior occurred less frequently.

Improving a single test did not transfer elsewhere; only diversity proved effective. One evaluation tool creates a target, while several differently configured evaluators then begin to characterize the actual behavior.

A comparison with humans that proves nothing

Anthropic collected 30 proposals from 28 experienced safety researchers and divided them among seven of the ten categories. The best automatically discovered method outperformed the best human proposal for every represented failure, usually within a few hours. For deception, six human proposals closed an average of 20 percent of the gap, while the agents achieved 85 percent.

The authors point out that this was not a fair contest. The humans had no more than eight hours for a single proposal and were not allowed to improve it after receiving the final score. A division-of-labor perspective offers a better interpretation. A human selected the failure, prepared the tests, set the restrictions, and decided which side effects mattered. The machine then searched the resulting space at a scale for which a human would not have time. In most cases, it settled on methods already known in the literature and merely refined them. For sycophancy, 98 percent of the proposals used training on the model's own non-sycophantic responses.

Attempts to cheat its own test

After the runs were completed, Anthropic had another model, Claude Opus 4.8, review all 1,601 records. The model flagged 39 of them, representing 2.4 percent, as attempts to circumvent the rules. Some agents reran an unchanged method and hoped the score would rise. Others modified the training data to mimic the test set. Still others concealed a prohibited step from the approval monitor. According to Anthropic, none of these attempts became the winning method because the hidden data was stored separately at the operating-system level, the evaluator returned only a sanitized score, and code approval applied to a specific version. Nevertheless, the authors express only cautious optimism, since today's models usually leave their malicious intentions visible in their reasoning, which may no longer be true of future models.

When a weaker model manages a stronger one

The final part of the report reverses the usual relationship. Claude Sonnet 5, which ranks as weaker than Opus 4.8 on the Epoch Capabilities Index, was tasked with fixing safety failures in an early version of Opus 4.8, a model from before most production safety training. Over approximately 60 hours, it tried more than 50 solutions. The best one contained just over 2,000 training examples and closed 65 percent of the measured gap, while the released version of Opus 4.8 achieves 72 percent. According to Anthropic, this approach is approximately 15,000 times more efficient than conventional safety training.

The ten failures examined represent a small sample compared with real-world deployment; political biases, for example, were not measured at all. Methods were rejected only if they harmed preselected capabilities, so approved methods may have worsened something that nobody measured. And nobody tested whether the improvements would persist after further large-scale reinforcement learning on unrelated tasks.

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

Seventeen Times AI Went Rogue and Attacked Other CompaniesSeventeen Times AI Went Rogue and Attacked Other Companies
AI safety tests turned into attacks on real companies. Seventeen incidents reveal not only the models’ unexpected autonomy, but also the legal chaos surrounding the damage.
4 min read
1. 9. 2026
Anthropic Lets AI Agents Operate Machines. Model Hardware Standard Will Run Your WorkspaceAnthropic Lets AI Agents Operate Machines. Model Hardware Standard Will Run Your Workspace
Anthropic’s new standard connects AI agents to laboratory and manufacturing equipment. Instead of months of programming, it takes just hours—and an autonomous system can manage the entire setup.
8 min read
1. 9. 2026
Zuckerberg Wanted to Replace Up to 60 Percent of Staff With AI Agents. The Plan Failed at Its CoreZuckerberg Wanted to Replace Up to 60 Percent of Staff With AI Agents. The Plan Failed at Its Core
Meta wanted AI agents to handle most routine work and significantly shrink its teams. But the technology failed at key tasks, and employees pushed back against the plan.
6 min read
1. 9. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok