A scorecard for the AI age
OpenAI ● Covered by 2 sources
OpenAI's CFO wants a new way to grade AI: not by hype, but by hard numbers. If it catches on, 'my chatbot is smart' stops being a business case.
Based on reporting by OpenAI — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Sarah Friar, OpenAI's chief financial officer, has never been shy about wanting AI spending held to the same standards as any other capital expense. Now she's putting a name to it: a scorecard, built around four questions that sound almost boringly practical for a company known for chasing artificial general intelligence.
The first metric is useful work — did the model actually finish the task, not just produce plausible-sounding text. The second is cost per successful task, which forces a comparison most vendors would rather avoid: what does it cost to get a correct answer, versus a wrong one that still burned compute. Third is dependability, essentially asking whether a system performs consistently enough that a business can build a process around it rather than babysit it. Fourth is return on compute, tying the whole exercise back to the one resource everyone in this industry is currently rationing.
What's notable here isn't the metrics themselves — engineers have quietly tracked versions of these for years. It's that OpenAI's own finance chief is saying them out loud, in public, as the industry's spending on GPUs, data centers and model training climbs into the hundreds of billions. Friar's framing implicitly challenges the current mode of AI marketing, where benchmark scores and demo videos stand in for actual evidence that a deployment pays for itself.
There's also a subtext of self-interest. OpenAI needs enterprise customers to keep renewing contracts, and renewals get easier when a company can point to a spreadsheet showing tasks completed per dollar rather than a vague sense that Copilot feels helpful. A scorecard like this gives sales teams and CFOs on the buyer's side a shared vocabulary, which is exactly what's been missing as AI budgets balloon without matching accountability.
My take — AI-written commentary, not fact-checked reporting
Finally, someone at a frontier lab is admitting that vibes aren't a KPI. I'd bet money this scorecard shows up in every enterprise AI pitch deck by next quarter, not because it's rigorous, but because CFOs everywhere are desperate for cover to justify the compute bill they already approved.
Read more about this at: OpenAI