AI agents fail at 76% of tasks

AI agents fail at 76% of tasks

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
26. 1. 2026
4 minutes reading
AI agents fail at 76% of tasks

Artificial intelligence promises a revolution in professional services—a team of experienced experts available 24/7 at a fraction of the usual salary. But the reality is different so far, as shown by the new APEX-Agents benchmark from Mercor.

What is the APEX-Agents benchmark?

APEX-Agents is the first benchmark to test whether AI agents can actually handle long-term, complex tasks in professional services. Unlike previous tests, it creates a real working environment with actual files and applications.

The benchmark was created by 256 experts from the Mercor platform—former consultants from BCG and McKinsey, investment bankers from Morgan Stanley and Citigroup, and corporate lawyers from Disney and other Fortune 500 companies. These professionals, with an average of 12.9 years of experience, created 480 tasks divided into 33 different "worlds"—complex project scenarios.

Each "world" represents a realistic project. For example, a team of consultants from the fictional company NorthPoint Strategy Partners is working for the client PureLife Wellness on a five-year expansion plan. They must map global demand, evaluate consumer trends, and identify the most promising markets. The experts created emails, spreadsheets, presentations, and other documents—just as they would for a real client. Each world contains an average of 166 files and nine applications with a total of 63 tools. The tasks are demanding—experienced professionals estimate that they take 1–2 hours to complete.

Results: Gemini 3 Flash leads, but the success rate is low

Mercor tested eight AI agents, with each agent performing every task eight times—for a total of 30,720 attempts. The results are sobering. Google DeepMind's Gemini 3 Flash performed best, with a Pass@1 score (first-attempt success rate) of 24.0%. It was closely followed by OpenAI's GPT-5.2 at 23.0%. Anthropic's Claude Opus 4.5 and Gemini 3 Pro ranked third and fourth, both with 18.4%.

First-attempt success rate
First-attempt success rate.

In practice, this means that if you give an AI agent a random task, the best model will complete it correctly only one time out of four. It will fail three times out of four. The two open-source models tested—GPT-OSS-120B and Kimi K2 Thinking—scored below 5%, significantly worse than the commercial models.

Success rates vary by type of work. The agents performed best on investment banking tasks, where GPT-5 and GPT-5.2 achieved 27.3%. The best score for consulting tasks was 22.7% (GPT-5.2), while for legal tasks it was 25.9% (Gemini 3 Flash).

Success rate by type of work.
Success rate by type of work. From left to right: Investment Banking Analyst, Management Consultant, and Corporate Lawyer.

Consistency is a major problem

When the agents had eight attempts at each task (Pass@8), the best model, GPT-5.2, achieved a 40.0% success rate—15 percentage points higher than with a single attempt. This shows that agents have the capabilities, but they are inconsistent. Even more interesting is the Pass^8 metric, which evaluates success across all eight attempts. The best model, Gemini 3 Flash, achieved only 13.4%. This means that even if an agent completes a task successfully once, there is no guarantee that it can do so again.

Gemini 3 Flash uses almost five times as many tokens as GPT-5.2 and 54% more steps. Although it is effective, it is not economical. At the other end of the spectrum, Kimi K2 Thinking uses an average of 92 steps and 1.6 million tokens per task, demonstrating that more resources do not always mean better results. All agents failed with a score of zero in at least 40% of attempts. Kimi K2 Thinking often became stuck in a loop and timed out in 29.8% of cases. Tasks requiring file creation were more difficult than those requiring a console response. Gemini 3 Flash performed best in both categories, but its score dropped by 4.9% on file-based tasks.

None of the instructions ever required files to be deleted. Nevertheless, GPT-5.2 deleted 21 files, Grok 4 deleted six, and Gemini 3 Flash deleted five. Claude Opus 4.5, GPT-5, and Kimi K2 Thinking did not delete any files.

To evaluate the outputs, the experts created a rubric with criteria. Each task has an average of 4.06 criteria. The Gemini 3 Flash model serves as the judge and achieved 98.5% accuracy when tested on a sample of 747 criteria.

Test evaluation

The APEX-Agents results show that AI agents have considerable room for improvement. The best agents achieve less than 25% on Pass@1 and no more than 40% on Pass@8.

Agents are capable of performing complex professional work, but they do so inconsistently and with a high failure rate. For businesses, this means that AI agents can be useful as assistants that help with parts of tasks, but they are not yet ready to fully replace experienced professionals. Mercor has made the entire dataset available as open source on Hugging Face, along with the Archipelago infrastructure for running and evaluating agents.

The future of work will probably not be about replacing people with AI, but about collaboration between people and AI agents, with each contributing their strengths. As AI agents improve, it will be interesting to see when—if ever—they reach a level at which they can reliably perform professional work without human supervision.

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

Altman Announced the Singularity Days After His Models Escaped the Lab on Their OwnAltman Announced the Singularity Days After His Models Escaped the Lab on Their Own
OpenAI chief Sam Altman declared on the Relentless podcast that humanity has already entered the singularity. “We’re like, in the singularity now,” he said verbatim. For decades, the term belonged more to science-fiction literature
6 min read
28. 7. 2026
AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.
Since April, Australian radio has been playing a dance remake of Madonna’s hit Like a Prayer on repeat. Released by Queensland DJ Josh Fawaz, it tops the radio airplay chart and has 35 million Spotify streams.
6 min read
28. 7. 2026
Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?
A first-person shooter that runs directly in the browser, with its own physics and eleven separate code modules. Around 55,000 lines in total, split across eleven subsystems and built on Thr
4 min read
28. 7. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok