Harmondale
Guide22 min

Why your AI spend does not turn into ROI

A field guide to why AI budgets rise without measurable value, and how to put every use case back under evidence.

Last updated: 25 June 2026

TLDR

  • 01

    The first problem is rarely the model. It is the missing line between usage, full cost, measured gain, and a business owner willing to be judged on that gain.

  • 02

    AI spend grows in layers: seats opened too quickly, pilots that never close, hidden API calls, human time spent correcting outputs, and governance added after the fact.

  • 03

    A serious AI ROI calculation happens use case by use case, with a control period or baseline, a quality measure, human review cost, and a stop rule if the signal remains weak.

  • 04

    The answer is not to cut AI everywhere. It is to move budget toward three to five workflows where AI removes a real bottleneck, then make every tool pay for itself with proof.

Symptom

The budget grows, but nobody can point to the result line that moved.

01

The common story starts cleanly. One team buys a few licenses to write faster, another connects an assistant to support, finance approves a budget because the risk of falling behind feels larger than the risk of waste, and the leadership team finally sees a modern direction. Six months later, usage looks impressive, screenshots circulate, employees say they save time, but the useful number is missing. Revenue has not moved, customer cycle time has not fallen in a verifiable way, margin has not improved, and the budget still renews.

The disconnect happens because the company mistakes a visible cost for an invisible transformation. A SaaS invoice is easy to approve, while real value depends on less glamorous changes: removing a step, changing a responsibility, closing an old tool, shortening an approval delay, or accepting that a task is now done differently. Until those changes are named, AI remains a layer placed on top of existing work. It adds local speed, then the organization absorbs that speed as rework, verification meetings, private prompts, and files nobody can audit.

An AI use case without a value owner becomes a comfort subscription. It may feel useful, but it is not governable.

Common error

Adoption metrics create a feeling of progress, not proof of return.

02

Adoption dashboards are tempting because they generate fast numbers: active users, prompts sent, documents generated, self-reported minutes saved. The problem is that these numbers describe tool activity, not enterprise effect. A team can send ten thousand prompts and simply move the work from writing to review. A support team can generate answers faster and increase escalations if the content lacks precision. A sales team can produce more email and damage reputation if the message becomes interchangeable.

Good measurement starts lower, inside the workflow. Which specific friction was supposed to disappear? How much did it cost before? What volume goes through that point? Which error becomes less frequent? Which delay gets shorter without quality dropping? Who confirms the gain is more than a feeling? These questions feel slow, but they protect against the most common trap: funding a machine that produces usage signals but constrains no decision. When usage rises and value does not follow, narrow the scope instead of demanding more adoption.

  1. 01

    Replace activity KPIs with one outcome metric, one quality metric, and one full-cost metric.

  2. 02

    Compare the equipped team against a period, workflow, or segment baseline before declaring a gain.

  3. 03

    Include human review time, because an unchecked AI output is not a gain; it is a liability.

Full cost

The real cost of an AI use case goes far beyond the license price.

03

A fifty-dollar seat looks modest until it is multiplied by dormant seats, redundant tools, workflows that need two extra validations, and IT teams that must secure behavior that spread without architecture. Full cost includes training, context loss when everyone invents a method, lost trust when a wrong answer reaches a customer, legal cost if sensitive data enters an unapproved tool, and the opportunity cost of the real problems the company did not address while celebrating adoption.

That is why public failures are useful. Zillow Offers was not simply a software issue: an algorithmic conviction touched the balance sheet, home inventory, and jobs. McDonald's did not merely test a voice interface: the drive-thru experiment met the real world, noise, accents, ambiguous orders, and customer tolerance. In both cases, the bill was not limited to the system. It spread to the operating model around the system. AI becomes expensive when it is treated as an isolated tool while it changes an entire decision chain.

Zillow

Zillow Offers was shut down after heavy losses connected to a home buying and selling operation that depended on forecasts difficult to control in a volatile market.

When a model influences committed capital, ROI must include downside exposure, not only average accuracy.

McDonald’s

The voice-ordering drive-thru test with IBM ended after mixed results and widely reported order mistakes.

A productivity gain does not exist if the operating environment creates too many exceptions for automation to stay fluid.

Decision

An AI use case needs a stop rule before it launches.

04

Most AI pilots are easy to start and hard to kill. Nobody wants to be the person slowing innovation, so the pilot becomes a permanent state. It continues because a few users like it, because it would be embarrassing to admit the promise was overstated, or because the invoice is small enough to avoid a dedicated meeting. That is how organizations create soft spend: too small to alarm, too scattered to optimize, too political to stop.

A stop rule changes the conversation. Before launch, the team writes the conditions that justify scaling, fixing, or ending the use case. For example: if handling time does not fall by 15 percent after six weeks without an error increase, stop it; if review time consumes more than 30 percent of the claimed time saved, reconfigure it; if fewer than 60 percent of outputs are usable without major rework, limit the use case. This discipline does not kill experimentation. It makes experimentation adult by turning enthusiasm into a testable hypothesis.

A pilot without a stop rule is not a pilot. It is spend waiting for renewal.

Field reality

Declared gains must survive re-entry into the real calendar.

05

Many AI gains are measured in an artificial moment: a demo, an isolated task, an internal benchmark, a week when the team pays attention because the project is watched. The return to the real calendar is harsher. Requests arrive incomplete, data is messy, the manager is unavailable, the customer replies with an edge case, the tool produces a convincing but wrong answer, and someone repairs the output. If measurement does not capture this friction, it sells a gain that exists only in the demo room.

The useful test follows the full path: initial request, data preparation, generation, verification, correction, approval, delivery, customer response, possible rework. AI can be excellent on one step and neutral or negative on the whole. It can also look modest locally and be very profitable if it removes a rare and expensive bottleneck. The question is not whether the tool is powerful. The question is whether the whole flow breathes better after it arrives.

Governance

Control is not a brake; it is what lets you increase budget with confidence.

06

Companies that oppose governance and innovation often end up with the worst of both: teams experimenting in the shadows and leaders unable to see what they fund. Useful governance does not start with a heavy committee. It starts with a short use-case register: tool, team, data touched, decision affected, monthly cost, owner, value metric, main risk, next review date. This register gives leadership a usable view without suffocating teams in abstract bureaucracy.

The NIST AI Risk Management Framework and ISO/IEC 42001 point to a simple idea: AI risk is managed through the lifecycle, not in a decorative charter. ROI works the same way. Map, measure, manage, review. Once that loop exists, it becomes easier to invest harder where evidence is strong. Control is not the enemy of ambition. It is the condition that lets the company say yes to a promising use case without funding every use case by default.

The practical governance test is whether a team can request more budget and receive a clear answer. If the answer depends on politics, excitement, or vendor pressure, the system is still weak. If the answer depends on measured value, remaining risk, quality threshold, and comparison with other use cases, governance is doing its job. It turns AI spend from a collection of enthusiastic exceptions into a portfolio that can be defended in front of finance, security, and the business.

  1. 01

    Create a use-case register with one business owner and one technical owner.

  2. 02

    Define the proof level required by financial, customer, legal, or human impact.

  3. 03

    Review expensive, sensitive, or fast-growing use cases monthly.

Resolution

The healthy path is a portfolio of use cases, not a collection of tools.

07

To escape AI spend without ROI, stop managing by vendor. A vendor sells capability; the company buys an outcome. The map should therefore start from workflows: support, sales, finance, operations, HR, legal, product, leadership. In each workflow, look for bottlenecks where volume is real, pre-AI cost is visible, quality can be checked, and human decision rights remain clear. Only then should the company choose the tool or keep the tool it already has.

The final portfolio should be deliberately small. Three well-instrumented use cases beat twenty-five pleasant experiments. Each use case gets a hypothesis, a full cost, an outcome metric, a quality metric, a stop rule, and a production plan. Savings from dormant seats and duplicate tools fund the use cases that prove something. It is less spectacular than a transformation announcement, but it is how AI becomes a management decision again instead of an act of faith.

Public failures

What visible cases teach ordinary deployments.

Public signal used as a reference point, not as a complete audit of the named company.

Zillow

Zillow shut down Zillow Offers after discovering that forecasting and real estate execution exposed the balance sheet to more volatility than the operating model could absorb.

When AI touches capital decisions, measure possible loss, not only average upside.

Deloitte Australia

An AI-assisted government report was corrected after problematic references and citations, with a partial refund reported.

Time saved in drafting is profitable only if verification, correction, and reputation costs stay below the gain.

Air Canada

A chatbot gave incorrect information about a commercial policy, and the company was held responsible for the answer on its website.

Customer automation should be measured by accuracy and accountability, not only ticket deflection.

Deployment

Deploy AI ROI control in 30 days

The goal is not to prove AI good or bad. The goal is to make every euro defensible, comparable, and reallocatable.

01

Inventory real use cases, not declared tools

Ask teams where AI actually enters their work: personal prompts, extensions, copilots, automations, data exports, meeting summaries, drafts sent to customers. For each use case, record monthly volume, data handled, direct cost, human time around the tool, and the decision that depends on the output.

ArtifactAI use-case register with owner, data, cost, volume, and risk.

02

Choose five workflows to measure first

Do not try to quantify everything at once. Select use cases where spend is high, risk is sensitive, or operational gain looks plausible. Each workflow needs a measurable before state: delay, cost, error, volume, satisfaction, margin, or rework rate.

ArtifactPrioritized workflow list with value hypothesis and comparison baseline.

03

Calculate full cost with review and exceptions

Add licenses, API calls, integration, training, supervision, correction, meeting time, escalations, prompt maintenance, and quality risks. This often reveals that the cheapest tool becomes expensive when it creates invisible rework.

ArtifactFull-cost model by use case with direct, human, and risk costs.

04

Install a stop rule and a scale rule

Define in advance what justifies stopping, correcting, or scaling. A use case that does not beat its value threshold after a defined period should be reduced or stopped. A use case that beats it can receive more budget only if quality and governance keep up.

ArtifactDecision sheet with value, quality, risk thresholds, and review date.

05

Reallocate budget instead of only reducing it

Successful rationalization is not cutting all AI spend. It cuts duplicates, dormant seats, and pilots without evidence to fund use cases that remove a real bottleneck. This is when leadership can say no to some tools and yes more strongly to a few workflows.

ArtifactReallocation plan with savings recovered and use cases strengthened.

FAQ

Should we stop AI licenses if ROI is not proven?

Not automatically. First separate useful but poorly measured use cases from genuinely weak ones. Cut dormant seats and duplicates quickly. For active workflows, give a short period with metrics, full cost, and a stop rule.

How do we measure time saved without relying on self-reporting?

Use a comparison baseline: same workflow before and after, equipped and non-equipped segments, or a reference period. Add a quality measure and rework time. A claimed gain that increases errors or approvals is not ROI.

Who should own AI ROI?

The business team that receives the value should own the hypothesis and the decision. IT secures, integrates, and observes. Finance helps with full cost. Without a business owner, the use case becomes a permanent experiment nobody can defend.

Diagnostic

Want to know where your AI is really leaking value?

We map your usage, hidden costs, and the points where AI should be cut, governed, or strengthened.

Diagnose my AI ROI