GDPval: OpenAI Tests How AI Handles Real-World Human Work

GDPval: OpenAI Tests How AI Handles Real-World Human Work

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
29. 9. 2025
3 minutes reading · 7 views
GDPval: OpenAI Tests How AI Handles Real-World Human Work

GDPval: OpenAI Tests How AI Handles Real Work and People

OpenAI introduced GDPval as a new tool for evaluating artificial intelligence performance on tasks with real economic value. This benchmark focuses on 44 occupations across nine major industries that contribute the most to the gross domestic product of the United States. These include healthcare, finance, manufacturing, and the government sector. The full set contains 1,320 tasks, while the open gold subset contains 220 tasks. Each task is based on the real work of professionals with an average of 14 years of experience, such as lawyers, engineers, and nurses.

The tasks are not just simple text-based questions. Models must work with various files, including spreadsheets, presentations, images, videos, and audio recordings. For example, in one manufacturing task, the model must design a device for testing cable reels while creating a PowerPoint presentation and a PDF summary. The average time for an expert to complete such a task is 7 hours, with an estimated value of around CZK 9,000.

How Did OpenAI Select the Occupations and Tasks?

The industries were selected based on data from the Federal Reserve Bank of St. Louis, with each contributing more than 5% to US GDP. Within each industry, OpenAI then selected the five predominantly digital occupations with the highest contribution to wages. To do so, it used the O*NET database from the US Department of Labor, classifying occupations as digital if at least 60% of their tasks were digital.

The occupations include software developers, lawyers, accountants, nurses, financial analysts, and editors. The experts who created the tasks underwent interviews, training, and reviews. Each task went through an average of five rounds of review, including automated checks by a model and evaluations by other professionals. This ensured that the tasks were representative and of high quality, with an average difficulty of 3.32 on a scale from 1 to 5.

Test Results: AI Is Approaching Human Performance

OpenAI tested models such as GPT-4o, o4-mini, o3, GPT-5, Claude Opus 4.1, Gemini 2.5 Pro, and Grok 4. The evaluation was conducted blindly, with experts comparing model outputs against human outputs. Claude Opus 4.1 achieved the best results, with its output rated as equal to or better than a human's in 47.6% of cases. GPT-5 excelled in accuracy, while Claude stood out in aesthetics, such as document formatting.

Results table

Model performance is improving linearly—from GPT-4o to GPT-5, it more than tripled. Models complete tasks up to 100 times faster and more cheaply than experts, with the average cost of a task for a human being around CZK 9,000. When combined with human oversight, such as review and corrections, models can save both time and money—for example, in a scenario where an expert tries the model several times and then completes the task themselves.

Model Errors and Improvements

The most common model issues were failures to follow instructions, formatting errors, and data hallucinations. For example, GPT-5 mainly failed on formatting, while Gemini and Grok often ignored references or promised outputs they did not deliver. Increasing reasoning effort, providing better context, or using scaffolding improved the results—for example, a special prompt for GPT-5 reduced PDF errors by half.

Table of model failures

OpenAI also experimented with less fully specified tasks, where models had less context. This led to worse results because they tried to infer the necessary information.

Open Resources and the Future

OpenAI released 220 tasks from the gold subset, including prompts and reference files, available at evals.openai.com. The site also includes an experimental automated evaluator that achieves 66% agreement with human experts. This tool is faster, but it has limitations, such as no internet access and issues with fonts.

GDPval has its limitations—it focuses on standalone digital tasks rather than interactive or physical work. Future versions are planned to expand to more occupations, interactivity, and more complex contexts. OpenAI is inviting experts and customers to contribute so that the benchmark can grow together with the community. This tool helps track AI progress and its impact on the labor market, where models can assist with routine work and leave the creative aspects to people.

Advertisement

Content created with help from UpTier.

SEO and GEO on autopilot. UpTier’s multi-agent systems write and optimize content for search engines and AI answers.

Discover UpTier ↗

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

OpenAI gives Codex reusable cloud workspaces accessible from any deviceOpenAI gives Codex reusable cloud workspaces accessible from any device
Codex gains reusable cloud development environments, alongside voice controls in its CLI, code reviews in the ChatGPT desktop app and cloud-based security tools.
2 min read
2. 10. 2026
Amazon releases Strands Decider 2B for AI workflow decisionsAmazon releases Strands Decider 2B for AI workflow decisions
Strands Decider 2B selects from predefined options and returns a confidence score. The fully open-source model is available now and small enough to run locally.
2 min read
1. 10. 2026
OpenAI says it disrupted a campaign to extract hidden model reasoningOpenAI says it disrupted a campaign to extract hidden model reasoning
OpenAI reported a coordinated effort to extract protected model reasoning and said it closed an extraction pathway. It attributed the main cluster of activity to individuals associated with Moonshot AI, the developer of Kimi.
3 min read
1. 10. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok