Sarah Friar, the CFO of OpenAI, published a scorecard for measuring AI value last week, and what I like most about it is the question it puts at the center.
How do we quantify the real value of AI?
That question has been badly served so far. We have counted licenses, logins, and pilot programs, and none of those numbers tell you whether anything of value got made.
Friar proposes a metric she calls "Useful Intelligence per Dollar," built on four questions:
- Is AI completing work that matters?
- What does each successful task cost?
- Can people depend on the result?
- Does each AI dollar produce more value as usage grows?
The elements of this framework together should look something like this:

Friar suggests scoring every AI output into three buckets: ready to use, needs correction, or needs escalation.
This is a better instrument than model accuracy, because it measures the work rather than the model. It shows you exactly where human labor is still going and whether that labor is shrinking.
Try it this week. Ask your team to sort their last ten AI tasks into those three buckets. An afternoon of that will teach you more than a quarterly survey.




