GDPval: OpenAI Tests How AI Handles Real Work and People
OpenAI introduced GDPval as a new tool for evaluating artificial intelligence performance on tasks with real economic value. This benchmark focuses on 44 occupations across nine major industries that contribute the most to the gross domestic product of the United States. These include healthcare, finance, manufacturing, and the government sector. The full set contains 1,320 tasks, while the open gold subset contains 220 tasks. Each task is based on the real work of professionals with an average of 14 years of experience, such as lawyers, engineers, and nurses.
The tasks are not just simple text-based questions. Models must work with various files, including spreadsheets, presentations, images, videos, and audio recordings. For example, in one manufacturing task, the model must design a device for testing cable reels while creating a PowerPoint presentation and a PDF summary. The average time for an expert to complete such a task is 7 hours, with an estimated value of around CZK 9,000.
How Did OpenAI Select the Occupations and Tasks?
The industries were selected based on data from the Federal Reserve Bank of St. Louis, with each contributing more than 5% to US GDP. Within each industry, OpenAI then selected the five predominantly digital occupations with the highest contribution to wages. To do so, it used the O*NET database from the US Department of Labor, classifying occupations as digital if at least 60% of their tasks were digital.
The occupations include software developers, lawyers, accountants, nurses, financial analysts, and editors. The experts who created the tasks underwent interviews, training, and reviews. Each task went through an average of five rounds of review, including automated checks by a model and evaluations by other professionals. This ensured that the tasks were representative and of high quality, with an average difficulty of 3.32 on a scale from 1 to 5.
Test Results: AI Is Approaching Human Performance
OpenAI tested models such as GPT-4o, o4-mini, o3, GPT-5, Claude Opus 4.1, Gemini 2.5 Pro, and Grok 4. The evaluation was conducted blindly, with experts comparing model outputs against human outputs. Claude Opus 4.1 achieved the best results, with its output rated as equal to or better than a human's in 47.6% of cases. GPT-5 excelled in accuracy, while Claude stood out in aesthetics, such as document formatting.

Model performance is improving linearly—from GPT-4o to GPT-5, it more than tripled. Models complete tasks up to 100 times faster and more cheaply than experts, with the average cost of a task for a human being around CZK 9,000. When combined with human oversight, such as review and corrections, models can save both time and money—for example, in a scenario where an expert tries the model several times and then completes the task themselves.
Model Errors and Improvements
The most common model issues were failures to follow instructions, formatting errors, and data hallucinations. For example, GPT-5 mainly failed on formatting, while Gemini and Grok often ignored references or promised outputs they did not deliver. Increasing reasoning effort, providing better context, or using scaffolding improved the results—for example, a special prompt for GPT-5 reduced PDF errors by half.

OpenAI also experimented with less fully specified tasks, where models had less context. This led to worse results because they tried to infer the necessary information.
Open Resources and the Future
OpenAI released 220 tasks from the gold subset, including prompts and reference files, available at evals.openai.com. The site also includes an experimental automated evaluator that achieves 66% agreement with human experts. This tool is faster, but it has limitations, such as no internet access and issues with fonts.
GDPval has its limitations—it focuses on standalone digital tasks rather than interactive or physical work. Future versions are planned to expand to more occupations, interactivity, and more complex contexts. OpenAI is inviting experts and customers to contribute so that the benchmark can grow together with the community. This tool helps track AI progress and its impact on the labor market, where models can assist with routine work and leave the creative aspects to people.



