Counting every generated word as completed work
Business measurement guide · Checked August 13, 2026
How to calculate AI ROI without making the spreadsheet lie
AI can save time, improve quality, create capacity, and introduce new costs. The calculation is useful only when it connects a defined workflow to a baseline, counts the work required to operate the system, and distinguishes measured financial return from an attractive hypothesis.
Michael Okeje
AI economics research and workflow measurement · Last updated August 13, 2026
ROI is a measurement system, not a magic multiplier
I am skeptical of AI business cases that begin with a universal claim such as “every employee saves two hours a day.” The claim may be directionally plausible for a particular task, but it is not a financial result. The team may use the time for higher-value work, or it may spend the time reviewing output. The tool may be used by only a fraction of employees. The benefit may be real but not appear as a reduction in payroll.
The better approach is to make the unit of analysis small. Measure one task, one group, one period, one baseline, and one outcome. Then expand the scope only when the first result is repeatable. McKinsey's 2026 measurement guidance makes the same essential connection: technical performance, user adoption, operational change, and financial impact must be linked with clear accountability. A model score alone cannot establish business value.
Step 1: write the calculation before the pilot
Write the expected value chain in plain language. For example: “The assistant drafts low-risk billing replies. A trained agent reviews them before sending. We expect lower handling time without increasing policy errors or escalations.” This statement tells you what to measure and what would invalidate the benefit.
“Adoption” belongs in the calculation because a perfect workflow used by nobody creates no operational return. “Eligible volume” prevents you from applying a result from a narrow task to every activity in a department. “Change per task” should come from observed cases, not a product promise. The formula is simple enough to audit and detailed enough to expose the assumptions.
Step 2: measure five layers of value
Do not force every benefit into a dollar amount. Report the layers separately, then state which ones support the financial calculation and which remain strategic or qualitative.
| Layer | Question | Evidence |
|---|---|---|
| Direct cost | Does the workflow require fewer paid hours, outsourced services, or avoidable operating expenses? | Verified labor, vendor, error, or processing cost change |
| Capacity | Can the same team complete more valuable work without lowering quality? | Completed work, backlog, throughput, and redeployed hours |
| Revenue | Does the workflow improve conversion, retention, speed to market, or a billable outcome? | Controlled comparison, attribution rules, and revenue evidence |
| Quality and risk | Are errors, delays, harmful actions, or compliance exposures reduced? | Incident rate, rework, escalation, loss avoided, expert assessment |
| Strategic option | Does the investment create a capability that opens a defensible path? | Milestones, customer evidence, differentiation, and future cost to exercise |
This separation matters because “strategic value” can be genuine while still being too uncertain to fund an immediate scale decision. State the uncertainty instead of converting it to a made-up revenue number.
Step 3: count the cost of making the output usable
The subscription is often the easiest cost to see and the easiest one to overemphasize. An AI workflow has a cost-to-serve. Include these lines in the period you are measuring:
If the workflow has model or tool usage, measure the distribution, not only the average. A long document, repeated tool call, difficult customer, or failed retrieval can create a high-cost tail. If the vendor bill cannot be capped or observed, include that uncertainty in the risk decision.
Worked example: an AI-assisted research brief
A 12-person US consulting team produces 80 internal research briefs each month. The current process takes 90 minutes per brief, including search, synthesis, and editing. A supervised AI workflow reduces first-draft time, but every brief still receives 25 minutes of analyst review. After a four-week pilot, 60% of eligible briefs use the workflow.
| Item | Calculation | Monthly result |
|---|---|---|
| Gross time reduction | 80 briefs x 60 minutes x 60% x $45 loaded hour | $2,160 potential capacity value |
| Incremental review | 80 x 25 minutes x 60% x $45 loaded hour | $900 review cost |
| Tool and usage | Subscription, search, storage, and model calls | $420 measured cost |
| Training and setup | Allocated monthly cost of initial setup and coaching | $300 in month one |
| Net measured benefit | $2,160 - $900 - $420 - $300 | $540 in month one |
If we called the entire 60-minute reduction “savings,” the result would look much larger. The more honest first-month calculation recognizes review, adoption, setup, and usage. It also asks whether the $1,260 of capacity value is actually redeployed to billable or otherwise valuable work. If not, report it as capacity released rather than cash saved.
Avoid five common ROI traps
Treating time released as payroll savings without a redeployment plan
Using a vendor's benchmark as if it were your own baseline
Ignoring review, rework, errors, support, and opportunity cost
Monetizing quality or strategic value with an invented dollar figure
A transparent negative or modest result is more valuable than inflated arithmetic. It tells you whether to improve the workflow, change the tool, reduce scope, or invest elsewhere.
A 90-day measurement loop
- Before launch: define the baseline, eligible volume, quality bar, adoption target, full cost, and decision threshold.
- Days 1-30: run a supervised pilot and record every accepted output, edit, failure, escalation, retry, and review minute.
- Days 31-60: compare user groups or time periods, inspect the high-cost and high-risk cases, and confirm whether released capacity is actually used.
- Days 61-90: test the workflow at the proposed scale, update the total-cost model, and document what changed from the pilot.
- At the decision: scale only if the outcome, quality, safety, adoption, economics, and operational owner all meet the agreed bar. Otherwise extend for a named question, redesign, or stop.
Use the site's AI pilot plan to structure the experiment, and AI agent business models to connect workflow economics to pricing.
Frequently asked questions
What is the formula for AI ROI?
A basic ROI formula is (financial benefit minus total investment) divided by total investment, multiplied by 100. For AI, the hard part is defining both terms honestly. Total investment includes software, model usage, implementation, integration, training, human review, support, governance, and opportunity cost. Benefits should be tied to measured changes in cost, revenue, quality, risk, or capacity rather than assumed from a fluent output.
What costs should be included in AI ROI?
Include subscriptions, API and model calls, storage, retrieval, tools, integration, engineering, configuration, evaluation, security, training, change management, human review, support, failed runs, rework, and the value of staff time assigned to the initiative. For a fair comparison, include recurring and one-time costs over the same period as the benefit.
How do you measure time saved by AI?
Measure the current time required for a defined task, the time with AI, the time spent reviewing or correcting output, and the number of eligible tasks actually using the workflow. Multiply only the time that is genuinely released or redirected to valuable work by a defensible loaded cost. Time that becomes extra capacity is not automatically cash savings.
What is a good AI ROI benchmark?
There is no universal benchmark that applies across workflows. A reported percentage can describe a survey population, a forecast, a self-reported benefit, or an audited financial result. Set a target from your baseline, risk, payback requirement, and alternative investment. Use external research as context, not as proof that your project will return the same amount.
How do I calculate AI ROI when the benefit is quality?
Translate quality changes into a measurable business consequence where possible. Measure errors, rework, escalations, refunds, delays, defects, retention, or expert review time before and after the workflow. If the benefit cannot be monetized credibly, report it separately as a quality or risk outcome instead of forcing a speculative dollar value into the ROI calculation.
When should an AI project be stopped?
Stop or redesign when the measured outcome does not improve after full review cost, the quality or safety threshold fails, adoption is too low to matter, data or contract conditions are unacceptable, the provider cannot support the workflow, or the opportunity cost exceeds the value. A negative result is useful if it prevents a larger investment in a poor use case.
Make the result auditable for someone who was not in the pilot
A finance or operations leader should be able to inspect the calculation without trusting the person who championed the tool. Keep the baseline cases, inclusion rules, usage logs, review sample, failure classification, and assumptions beside the result. Label estimates, observations, and forecasts separately. A number that cannot be reproduced is a talking point, not a measurement.
Keep a counterfactual in view. What would the team have done without the AI workflow? Would it have hired, outsourced, delayed, or simply accepted the same workload? The answer changes the value of time released. It also helps prevent the common mistake of claiming the full cost of an avoided hire when the organization never planned to hire anyone.
Review the result by segment. A workflow may work well for short documents but fail on long ones, improve experienced users but slow beginners, or help one region while introducing a language or policy problem in another. Report the slices that matter to the decision. An overall average should be a summary of the evidence, not a replacement for it.
Keep these records
Baseline definition, task volume, adoption, accepted and rejected outputs, review minutes, cost ledger, quality rubric, incidents, and decision notes.
Ask these questions
What changed? For whom? Compared with what? At what cost? With what risk? What would make the result stop being true?