Measuring the performance of our models on real-world tasks
OpenAI
OpenAI built a new test called GDPval to see how AI models handle real jobs, not just quizzes. It grades them on actual work across 44 professions, not trivia.
Based on reporting by OpenAI — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
OpenAI has a new benchmark, and this one's aimed at something more useful than chess puzzles or SAT questions. It's called GDPval, and the idea is simple: stop asking models to solve abstract brainteasers and start asking them to do the kind of work people actually get paid for.
The evaluation spans 44 occupations, which is a deliberate signal. OpenAI isn't just checking if a model can write code or summarize a document in isolation. It's trying to capture something closer to economic value, the tasks that show up in job descriptions and performance reviews rather than academic test banks. That's a meaningfully different bar than most existing benchmarks clear.
This matters because the industry has a measurement problem. Every few months a lab announces a new state-of-the-art score on some reasoning or coding leaderboard, and the numbers climb, and yet nobody outside the lab can say with confidence whether that translates into a model being genuinely helpful at, say, drafting a legal brief or handling a customer escalation. GDPval is an attempt to close that gap, at least for OpenAI's own models, by grounding evaluation in tasks tied to real occupations rather than synthetic puzzles.
It's also a tell about where OpenAI thinks the competition is heading. Benchmarks are marketing as much as they are science, and by naming this one after GDP, OpenAI is making an explicit claim: our models aren't just clever, they're economically productive. Whether that claim holds up under scrutiny from outside researchers is a different question. But the framing itself says a lot about what the company wants people to measure next.
My take — AI-written commentary, not fact-checked reporting
I'll believe an AI benchmark is meaningful when it's run by someone with nothing to gain from the score, and OpenAI grading its own homework on "economic value" is exactly the kind of self-serving metric that should make you squint. Useful concept, wrong referee — someone independent needs to run GDPval-style tests before anyone treats the results as more than a press release.
Read more about this at: OpenAI