AI Pulse by Inblix

OpenAI's GDPval tests if AI can do your job across 44 careers

OpenAI Blog · Jul 12, 2026 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: OpenAI's GDPval tests if AI can do your job across 44 careers

OpenAI is tired of AI benchmarks that feel like glorified SAT prep. The company just dropped GDPval, a new evaluation that tosses out synthetic exam questions and instead measures how well models handle the actual deliverables of 44 real-world occupations. We’re talking legal briefs, engineering blueprints, nursing care plans, and customer support logs—not multiple-choice trivia. The name isn’t subtle: it’s a direct nod to Gross Domestic Product, built by pulling tasks from the top nine industries contributing to U.S. GDP, like healthcare and professional services.

This isn’t another academic leaderboard flex. GDPval marks a deliberate step in OpenAI’s progression from sterile benchmarks like MMLU toward messier, market-based evaluations like SWE-Lancer. The dataset comprises 1,320 specialized tasks, with 220 being open-sourced in a “gold” set. Each task was vetted by professionals averaging over 14 years of experience, ensuring the work isn’t theoretical. Unlike simple text prompts, models have to wrestle with reference files, spreadsheets, diagrams, and multimedia to produce finished work products.

There’s a clear philosophical bet here. OpenAI argues that major technologies take a decade to diffuse, and evaluations like this ground the AGI conversation in current capability rather than sci-fi speculation. The company wants to track how AI might actually support people in their daily work, not just how it scores on a test.

But let’s not get carried away. This is version one, and it’s explicitly limited to one-shot evaluations. It doesn’t capture the iterative back-and-forth, the context-building, or the painful revision cycles that define most knowledge work. It’s a snapshot of a single perfect delivery, not a simulation of a chaotic Tuesday. The 44 occupations are a solid start, but the economy is a lot bigger. Right now, GDPval shows us what an AI can produce when the brief is perfect. It doesn’t yet show us if it can survive the meeting before the brief is written.

💡 Key Takeaways

  1. GDPval tasks are based on real work products like legal briefs and engineering blueprints, not synthetic exam questions, and were created by professionals averaging 14+ years of experience.
  2. The evaluation spans 1,320 tasks across 44 occupations in 9 industries, with 220 tasks being open-sourced, and requires models to handle multimedia, spreadsheets, and reference files.
  3. As a one-shot evaluation, GDPval currently cannot measure a model's ability to iterate, build context over time, or handle the interactive workflows common in real jobs.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles