GDPval: OpenAI Tests How AI Handles Real-World Human Work

GDPval: OpenAI Tests How AI Handles Real-World Human Work

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
29. 9. 2025
3 minutes reading
GDPval: OpenAI Tests How AI Handles Real-World Human Work

GDPval: OpenAI Tests How AI Handles Real Work and People

OpenAI introduced GDPval as a new tool for evaluating artificial intelligence performance on tasks with real economic value. This benchmark focuses on 44 occupations across nine major industries that contribute the most to the gross domestic product of the United States. These include healthcare, finance, manufacturing, and the government sector. The full set contains 1,320 tasks, while the open gold subset contains 220 tasks. Each task is based on the real work of professionals with an average of 14 years of experience, such as lawyers, engineers, and nurses.

The tasks are not just simple text-based questions. Models must work with various files, including spreadsheets, presentations, images, videos, and audio recordings. For example, in one manufacturing task, the model must design a device for testing cable reels while creating a PowerPoint presentation and a PDF summary. The average time for an expert to complete such a task is 7 hours, with an estimated value of around CZK 9,000.

How Did OpenAI Select the Occupations and Tasks?

The industries were selected based on data from the Federal Reserve Bank of St. Louis, with each contributing more than 5% to US GDP. Within each industry, OpenAI then selected the five predominantly digital occupations with the highest contribution to wages. To do so, it used the O*NET database from the US Department of Labor, classifying occupations as digital if at least 60% of their tasks were digital.

The occupations include software developers, lawyers, accountants, nurses, financial analysts, and editors. The experts who created the tasks underwent interviews, training, and reviews. Each task went through an average of five rounds of review, including automated checks by a model and evaluations by other professionals. This ensured that the tasks were representative and of high quality, with an average difficulty of 3.32 on a scale from 1 to 5.

Test Results: AI Is Approaching Human Performance

OpenAI tested models such as GPT-4o, o4-mini, o3, GPT-5, Claude Opus 4.1, Gemini 2.5 Pro, and Grok 4. The evaluation was conducted blindly, with experts comparing model outputs against human outputs. Claude Opus 4.1 achieved the best results, with its output rated as equal to or better than a human's in 47.6% of cases. GPT-5 excelled in accuracy, while Claude stood out in aesthetics, such as document formatting.

Results table

Model performance is improving linearly—from GPT-4o to GPT-5, it more than tripled. Models complete tasks up to 100 times faster and more cheaply than experts, with the average cost of a task for a human being around CZK 9,000. When combined with human oversight, such as review and corrections, models can save both time and money—for example, in a scenario where an expert tries the model several times and then completes the task themselves.

Model Errors and Improvements

The most common model issues were failures to follow instructions, formatting errors, and data hallucinations. For example, GPT-5 mainly failed on formatting, while Gemini and Grok often ignored references or promised outputs they did not deliver. Increasing reasoning effort, providing better context, or using scaffolding improved the results—for example, a special prompt for GPT-5 reduced PDF errors by half.

Table of model failures

OpenAI also experimented with less fully specified tasks, where models had less context. This led to worse results because they tried to infer the necessary information.

Open Resources and the Future

OpenAI released 220 tasks from the gold subset, including prompts and reference files, available at evals.openai.com. The site also includes an experimental automated evaluator that achieves 66% agreement with human experts. This tool is faster, but it has limitations, such as no internet access and issues with fonts.

GDPval has its limitations—it focuses on standalone digital tasks rather than interactive or physical work. Future versions are planned to expand to more occupations, interactivity, and more complex contexts. OpenAI is inviting experts and customers to contribute so that the benchmark can grow together with the community. This tool helps track AI progress and its impact on the labor market, where models can assist with routine work and leave the creative aspects to people.

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

Altman Announced the Singularity Days After His Models Escaped the Lab on Their OwnAltman Announced the Singularity Days After His Models Escaped the Lab on Their Own
OpenAI chief Sam Altman declared on the Relentless podcast that humanity has already entered the singularity. “We’re like, in the singularity now,” he said verbatim. For decades, the term belonged more to science-fiction literature
6 min read
28. 7. 2026
AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.
Since April, Australian radio has been playing a dance remake of Madonna’s hit Like a Prayer on repeat. Released by Queensland DJ Josh Fawaz, it tops the radio airplay chart and has 35 million Spotify streams.
6 min read
28. 7. 2026
Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?
A first-person shooter that runs directly in the browser, with its own physics and eleven separate code modules. Around 55,000 lines in total, split across eleven subsystems and built on Thr
4 min read
28. 7. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok