Artificial intelligence promises a revolution in professional services—a team of experienced experts available 24/7 at a fraction of the usual salary. But the reality is different so far, as shown by the new APEX-Agents benchmark from Mercor.
What is the APEX-Agents benchmark?
APEX-Agents is the first benchmark to test whether AI agents can actually handle long-term, complex tasks in professional services. Unlike previous tests, it creates a real working environment with actual files and applications.
The benchmark was created by 256 experts from the Mercor platform—former consultants from BCG and McKinsey, investment bankers from Morgan Stanley and Citigroup, and corporate lawyers from Disney and other Fortune 500 companies. These professionals, with an average of 12.9 years of experience, created 480 tasks divided into 33 different "worlds"—complex project scenarios.
Each "world" represents a realistic project. For example, a team of consultants from the fictional company NorthPoint Strategy Partners is working for the client PureLife Wellness on a five-year expansion plan. They must map global demand, evaluate consumer trends, and identify the most promising markets. The experts created emails, spreadsheets, presentations, and other documents—just as they would for a real client. Each world contains an average of 166 files and nine applications with a total of 63 tools. The tasks are demanding—experienced professionals estimate that they take 1–2 hours to complete.
Results: Gemini 3 Flash leads, but the success rate is low
Mercor tested eight AI agents, with each agent performing every task eight times—for a total of 30,720 attempts. The results are sobering. Google DeepMind's Gemini 3 Flash performed best, with a Pass@1 score (first-attempt success rate) of 24.0%. It was closely followed by OpenAI's GPT-5.2 at 23.0%. Anthropic's Claude Opus 4.5 and Gemini 3 Pro ranked third and fourth, both with 18.4%.
In practice, this means that if you give an AI agent a random task, the best model will complete it correctly only one time out of four. It will fail three times out of four. The two open-source models tested—GPT-OSS-120B and Kimi K2 Thinking—scored below 5%, significantly worse than the commercial models.
Success rates vary by type of work. The agents performed best on investment banking tasks, where GPT-5 and GPT-5.2 achieved 27.3%. The best score for consulting tasks was 22.7% (GPT-5.2), while for legal tasks it was 25.9% (Gemini 3 Flash).
Consistency is a major problem
When the agents had eight attempts at each task (Pass@8), the best model, GPT-5.2, achieved a 40.0% success rate—15 percentage points higher than with a single attempt. This shows that agents have the capabilities, but they are inconsistent. Even more interesting is the Pass^8 metric, which evaluates success across all eight attempts. The best model, Gemini 3 Flash, achieved only 13.4%. This means that even if an agent completes a task successfully once, there is no guarantee that it can do so again.
Gemini 3 Flash uses almost five times as many tokens as GPT-5.2 and 54% more steps. Although it is effective, it is not economical. At the other end of the spectrum, Kimi K2 Thinking uses an average of 92 steps and 1.6 million tokens per task, demonstrating that more resources do not always mean better results. All agents failed with a score of zero in at least 40% of attempts. Kimi K2 Thinking often became stuck in a loop and timed out in 29.8% of cases. Tasks requiring file creation were more difficult than those requiring a console response. Gemini 3 Flash performed best in both categories, but its score dropped by 4.9% on file-based tasks.
None of the instructions ever required files to be deleted. Nevertheless, GPT-5.2 deleted 21 files, Grok 4 deleted six, and Gemini 3 Flash deleted five. Claude Opus 4.5, GPT-5, and Kimi K2 Thinking did not delete any files.
To evaluate the outputs, the experts created a rubric with criteria. Each task has an average of 4.06 criteria. The Gemini 3 Flash model serves as the judge and achieved 98.5% accuracy when tested on a sample of 747 criteria.
Test evaluation
The APEX-Agents results show that AI agents have considerable room for improvement. The best agents achieve less than 25% on Pass@1 and no more than 40% on Pass@8.
Agents are capable of performing complex professional work, but they do so inconsistently and with a high failure rate. For businesses, this means that AI agents can be useful as assistants that help with parts of tasks, but they are not yet ready to fully replace experienced professionals. Mercor has made the entire dataset available as open source on Hugging Face, along with the Archipelago infrastructure for running and evaluating agents.
The future of work will probably not be about replacing people with AI, but about collaboration between people and AI agents, with each contributing their strengths. As AI agents improve, it will be interesting to see when—if ever—they reach a level at which they can reliably perform professional work without human supervision.



