July 28, 2026

OpenAI CFO: Measure "Useful Intelligence per Dollar"

OpenAI CFO Sarah Friar proposes measuring AI by useful work, not adoption. Here is why the framework holds up, and what it means for your own AI fluency.
Daan van Rossum
By
Daan van Rossum
Founder & CEO, Lead with AI

Presented by

Sarah Friar, the CFO of OpenAI, published a scorecard for measuring AI value last week, and what I like most about it is the question it puts at the center.

How do we quantify the real value of AI?

That question has been badly served so far. We have counted licenses, logins, and pilot programs, and none of those numbers tell you whether anything of value got made.

Friar proposes a metric she calls "Useful Intelligence per Dollar," built on four questions:

  • Is AI completing work that matters?
  • What does each successful task cost?
  • Can people depend on the result?
  • Does each AI dollar produce more value as usage grows?

The elements of this framework together should look something like this:

AI - Useful Intelligence per Dollar

Friar suggests scoring every AI output into three buckets: ready to use, needs correction, or needs escalation.

This is a better instrument than model accuracy, because it measures the work rather than the model. It shows you exactly where human labor is still going and whether that labor is shrinking.

Try it this week. Ask your team to sort their last ten AI tasks into those three buckets. An afternoon of that will teach you more than a quarterly survey.

Flagship AI Newsletter
The AI Newsletter That Makes You Smarter, Not Busier
Join over 30,000 leaders and receive our insights on AI platforms, implementations, and organizational change management.
FlexOS Course - AI Content Accelerator - Testimonial Badge

Useful Work Is the Only Thing Worth Counting

The first question Friar posits, is AI completing work that matters, is the one that matters most to me.

Her advice is to start with the work itself: how many customer issues did AI help resolve? How many code changes did it help ship? How much time did it give back to people?

This lands directly on something I have argued for years. The only real benefit of AI is that it increases a person's impact per hour. Everything else, including adoption rates, tool counts, and enthusiasm in the all-hands, is a proxy at best and a distraction at worst.

So it is useful to see a CFO put work accomplished where seats purchased used to sit.

This Should Force the End of Loose Experimenting

Here is why the framing matters beyond measurement.

Most organizations are still experimenting with AI loosely. They try things, they buy things, they run pilots, and nobody asks the harder question of where the investment is actually worth making.

The cost of that habit is measurable. Glean's Work AI Index 2026 found that 37% of the time workers spend with AI goes to supervising and fixing its output, against 36% actually producing work with it, which comes to 6.4 hours a week. We covered this in Lead with AI's June Executive Briefing.

A scorecard changes that conversation. Once you are measuring cost per successful task rather than cost per license, you have to be deliberate about which tasks, and that discipline is the point. Our GED-RT framework exists for exactly this decision, because scoring your workflows before you invest is what separates a program from a pile of experiments.

Friar's math is simple. Add the full cost of completing the work, including employee time, human review, retries, and rework. Count the tasks that met your quality bar. Divide.

She also makes a point worth underlining twice: the lowest price per token does not produce the lowest cost per outcome. A more capable model can be better value if it gets the answer right in one pass, cutting retries, review time, and rework.

The Same Discipline Applies to Your Personal AI Fluency

Does this only work at the enterprise level? No, and this is where I think it gets most interesting.

The same question applies to how you invest in your own AI fluency. There is a real ladder here, and each rung costs more time than the last.

At the bottom, you write a better prompt, which is where our CODO prompting framework does most of its work by making you specify Character, Objective, Dos and Donts, and Outputs.

Above that, you customize an AI with your own context, your own examples, and your own standards. Higher still, you write a skill that encodes how a particular piece of your work should be done, so you never have to re-explain it. At the top, you build an actual AI app for a workflow that runs often enough to deserve one.

Every rung is a real investment of your time. So which rung is a given task worth? Ask how often you do the work, and how much of your judgment it needs each time. A monthly report that always follows the same shape deserves a skill. A one-off analysis deserves a good prompt and nothing more.

For the middle rungs, where you are building a small assistant for one repeated task, the OVER framework is the method we teach.

That is cost per successful task applied to yourself, and it is the difference between using AI and building with it.

The Most Practical Idea in the Piece

Two things are worth adding to the scorecard.

First, no scorecard will save an untouched workflow. McKinsey's July 2026 research found that leaders who redesigned their workflows were 5.3 times more likely to report real enterprise value than those who left them in place, at 32% against 6%. The same study found highly AI-fluent leadership teams were 3.9 times more likely to capture value than low-fluency ones. (Source: McKinsey, From Adoption to Impact, July 2026. More figures like these live in our running collection of AI statistics.)

Measurement is an instrument, not a strategy, and the redesign has to come first. Stanford researchers gave this capability a name, Workflow Literacy, and it is the thing that decides whether your scorecard has anything to report.

Second, Friar takes safe use seriously. Before AI moves from drafting to taking action, she says organizations should define what data it can access, what systems it can use or change, and when a person must review or approve. Those three questions are a practical governance starting point.

One caveat belongs here. Friar works for OpenAI, and the scorecard arrives with a recommendation to buy frontier models and ChatGPT Work. The four questions hold regardless of whose model you run, so keep the framework and skip the shopping list.

The Bottom Line

  1. Count useful work, not adoption. Impact per hour is the only number that ultimately matters.
  2. Set the North Star first. Time given back is only value if you decided what it was for.
  3. Stop experimenting loosely. Decide where AI investment is genuinely worth making, and measure that.
  4. Calculate cost per successful task, including review, retries, and rework.
  5. Apply the same test to your own fluency. Match the investment, from prompt to custom setup to skill to app, to how often the work repeats.
  6. Score outputs as ready to use, needs correction, or needs escalation, and watch the mix move.
  7. Redesign the workflow first, then measure it.

This is the work we do with leaders in AI Leader Advanced, where the whole point is deciding which workflows deserve your investment before you spend it.