The short version: OpenAI’s chief financial officer, Sarah Friar, published a scorecard on July 17 for measuring whether AI spending is actually working. She calls the core metric “Useful Intelligence per Dollar,” built from four parts: useful work completed, cost per successful task, dependability, and value at scale. It is aimed at enterprise finance teams, but the logic scales down cleanly. A small business owner can run a lightweight version of this same scorecard on a single AI tool this month, with a notebook and about twenty minutes.
What did OpenAI’s CFO actually propose?
In a blog post published on OpenAI’s site, Friar framed what she called “the basic economic question facing CFOs” right now: does the value of the work AI completes grow faster than the cost of producing it? Her answer is a four-part framework, reported in detail by Axios and CFO Dive.
The four measures: useful work completed (what the AI actually produced, whether that is customer issues resolved, contracts reviewed, or hours handed back to people), cost per successful task (the full cost, including human review and rework, divided by the number of tasks that met the quality bar, not the number attempted), dependability (how much output can be trusted without a person double-checking it), and value at scale (whether each dollar and each unit of compute produces more over time, not less).
Notice what is missing from that list: how many people are using the tool, how many prompts got sent, how much the subscription costs. Those are the numbers most businesses actually track. Friar’s scorecard is built almost entirely around outcomes instead.
Why does “cost per successful task” matter more than your subscription bill?
This is the part of the framework worth sitting with. A $20-a-month AI subscription looks cheap on paper. But if your team spends fifteen minutes fixing what it produced, the real cost per completed task can be considerably higher, once you count the salary time spent on the fix. Friar’s framework insists on dividing total cost, including that rework time, by the number of tasks that came out right the first time.
That is not a new idea on its own. We covered the enterprise version of this problem back in June, when Uber burned through an entire year’s AI budget in four months and companies started discovering that unmonitored AI spend does not automatically pay for itself. What is different here is that the company selling the AI tools is now the one publishing a formal method for catching that exact problem before it happens. That is a notable admission for OpenAI to volunteer.
How reliable does AI have to be before it is worth using?
Dependability is the least intuitive of Friar’s four measures, and possibly the most useful for a small business. The idea is simple: when AI output is accurate, well-sourced, and consistent enough that people stop double-checking it, review time drops and the tool becomes safe to use in bigger, more important workflows. When it is not, every task quietly costs more than it looks like it does, because someone has to verify it.
In practice, this means tracking something most business owners never write down: how often do you accept the AI’s first draft as-is, versus how often do you rewrite it? If you are correcting more than a third of what a tool produces for a given task, that tool has not earned its way into that workflow yet, regardless of what the subscription costs.
How can a small business run this scorecard without a finance team?
Friar built this for organizations with dedicated finance staff and usage dashboards. Most small businesses have neither. But the four questions translate into something you can run with a notebook.
Pick one task you already use AI for regularly: drafting client emails, summarizing meeting notes, writing product descriptions. For two weeks, log four things every time you use it: what you asked for, whether you used the output as-is or rewrote it, roughly how long the rewrite took if you needed one, and what the task would have cost you in time without the tool at all. At the end of two weeks, add up total time spent (using the tool plus fixing its output) and compare it against the time the task would have taken you unassisted. That ratio is your cost per successful task, in miniature.
Do this for one workflow before expanding to five. Our reporting earlier this year found that 87 percent of small business owners use AI daily, but only one in five feel confident enough to call it a revenue driver. A big part of that confidence gap is simply never having measured whether a given tool earns its keep. Friar’s framework, shrunk down, is a fast way to close that gap on your own terms rather than guessing.
What should you actually do with this?
If you are already juggling several AI subscriptions, run the two-week log on whichever one you are least sure about first. If it clears the bar, keep it and stop worrying about it. If it does not, that is real information, not a vague feeling that “the AI thing isn’t really working out.” For anyone still choosing between tools, our roundup of AI tools organized by the job you need done is a reasonable place to start narrowing options before you run the scorecard on the finalists.
The honest takeaway is that OpenAI did not invent a complicated new metric here. It gave a name and a structure to something disciplined operators already do instinctively: check whether a tool is actually paying for itself before assuming it is. Naming it clearly is still useful, because a framework you can write down is a framework you will actually use twice.
Frequently Asked Questions
What is OpenAI’s “Useful Intelligence per Dollar” scorecard?
It is a measurement framework published by OpenAI CFO Sarah Friar on July 17, 2026, built around four factors: useful work completed, cost per successful task, dependability, and value at scale. It is designed to help organizations judge whether AI spending is producing more value than it costs, rather than just tracking adoption or subscription numbers.
Can a small business actually use a framework built for enterprise CFOs?
Yes, in a scaled-down form. The underlying questions, what did the AI produce, what did it really cost including rework time, how much can you trust it without checking, and is it getting more efficient over time, apply at any business size. A small business can track this with a simple log on one workflow rather than a dashboard.
What is “cost per successful task” and why does it matter more than a subscription price?
It is the total cost of using an AI tool, including the time spent reviewing and correcting its output, divided by the number of tasks that met your quality bar on the first try. A cheap subscription can still produce an expensive cost per successful task if your team spends significant time fixing what it generates.
How do I know if an AI tool is “dependable” enough to keep using?
Track how often you accept its output without changes versus how often you need to rewrite or correct it for a given task. If you are reworking more than roughly a third of what it produces, the tool has not yet earned its place in that specific workflow, even if it might be a good fit for a different one.
Have you ever actually sat down and measured whether a specific AI subscription is paying for itself, or is it still running on a gut feeling that it probably helps? What would you find if you tracked it for two weeks?
