Just Launched Gruve PulseAI Platform, your private AI infrastructure, production-ready in under 2 weeks.PulseAI is live — private AI, ready in 2 weeks.

See PulseAI
Blog

The denominator problem: Why enterprise AI ROI math doesn’t work yet, and what to do about it before your next budget cycle

August 6, 2026

Every board conversation I am in right now eventually arrives at the same question. What are we getting for this?

It is the right question. It is also, at this moment, close to unanswerable, and not for the reason most people assume. The industry has spent a year arguing about the return. The harder problem is the investment.

I was talking through this recently with Toby Stuart, who teaches strategy at Berkeley Haas, sits on a number of late stage technology boards, and wrote the Harvard Business Review piece “The Future Is Shrouded in an AI Fog.” He put the problem more plainly than I had heard it put before.


Toby Stuart, Berkeley Haas

That is the whole thing. We keep debating the numerator while the denominator is still fuzzy. And you cannot run the math until you can state both.

You do not actually know what you are spending

Ask a technology leader what their organization spends on AI and you will get a number. It will be wrong, and usually by a lot.

The number people quote is the visible line. The platform contract, the model provider bill, the bespoke agent deployed against a specific workflow. But every piece of enterprise software in your estate now has AI embedded in it, and your people use it constantly without ever thinking of it as AI spend. Your SaaS vendors are repackaging that consumption as token credits and selling it back to you inside renewals nobody is reading line by line. The spend is distributed, behavioral, and largely unattributed.

Then there is the part that is visible but not controlled. Uber burned through its entire 2026 AI budget in four months after rolling Claude Code out to roughly 5,000 engineers. Not misuse. Engineers used the tool for exactly what it was built for. One executive ran up $1,200 in a single two hour session. The company has since capped tooling at $1,500 per engineer per month, and its president and COO, Andrew Macdonald, said the link between that consumption and anything a customer would notice “is not there yet.”

I have said before that companies running on demo scale assumptions start losing good money overnight once agents run continuously. Uber is not a cautionary tale about recklessness. It is a preview of what happens to any organization that scales consumption faster than it scales attribution.

Token maxing was a phase, not a strategy

For most of the past year the prevailing advice was to maximize token consumption. Get people using it. Drive experimentation. There was a defensible reason for that: compute was constrained, adoption was the bottleneck, and the fastest way to find real use cases was to lower the friction to trying things.

But Toby framed the flaw in it as a basic leadership principle.

That phase served its purpose. It should end now, because the ground under inference economics has shifted. Kimi K3 shipped open weights on July 27 at 2.8 trillion parameters. Chinese models now account for roughly 60 percent of token usage by U.S. companies on OpenRouter. The routing layer between your application and the intelligence is becoming a real engineering discipline with real margin attached to it, which means unoptimized spend is no longer just expensive. It is a competitive gap.

This is why FinOps and governance are the same conversation, and why both are the shift of the next 18 months. Not because governance is a brake. Because you cannot optimize what you cannot attribute, and you cannot attribute what you did not instrument.

The return side is worse

The spend side is at least tractable. Instrument it, attribute it, route it. Hard, but bounded.

The return side is where this gets genuinely difficult, and the reason is not measurement error. It is that in most organizations the return has not been captured yet, because the work around it never changed.

Toby put it this way.

The data backs that up. Roughly 48 percent of organizations report introducing AI without redesigning the workflow or the role it sits inside. Only about 12 percent report redesign at scale with a new operating model behind it. Fewer than one in three CFOs can point to a specific financial return.

If you hand a team four hours back per person per week and change nothing else about how that team is structured, measured, or staffed, you have not created value. You have created slack. Slack is not a bad thing. It is just not something that shows up on a P&L, and it is not what you told the board you were buying.

There is also a category of benefit that resists measurement entirely. Take the oldest problem in any large distributed organization: you do not know who knows what, so you constantly reinvent the wheel. General intelligence applied across a company’s own work product should solve a meaningful share of that. But how do you price the week someone did not spend rebuilding something a colleague had already built? There is no standard metric for a connection that would not otherwise have happened. Outside of narrow deployments with clean before and after cost measures, as Toby puts it, measurement remains deeply problematic.

Two waves, not one

The way I think about the timing is seismic. A P wave and an S wave.

The P wave is the first pop. It arrives fast, it is visible, and it is what most organizations are currently measuring: a workflow that got quicker, a support queue that got shorter, a tier one triage load that fell. It is real value and you should take it.

The S wave is the second arrival, and it is the larger one. It comes from the retooling. Reorganized teams, redesigned roles, reallocated resources, a shape of company that is actually built for how the work now gets done. That wave lands later and it is much harder to attribute cleanly to any single investment, which is exactly why so many ROI models miss it.

The mistake is running a three year NPV model on the P wave and concluding the technology did not pay. The other mistake is promising the S wave on a P wave timeline.

Instrument first, then optimize, then argue about ROI

Here is where I have landed, and it is what we are doing inside Gruve with our own P&L.

We give engineers quota budgets to build against real opportunities. We capture what that costs. We have accounting report it back into our own intelligence platform so consumption is attributable to the work that generated it. We are asking the same question of our cloud accounts and our token consumption platforms.

We are not getting ROI in full scope from that. Nobody is. But we are getting visibility, and that is where it starts. You cannot govern what you cannot see, you cannot optimize what you cannot attribute, and you cannot defend a number to a board that you cannot decompose.

Two things are worth holding onto while the fog is this thick.

Liability does not travel with the workload. You can run inference on someone else’s infrastructure, but if an AI driven decision goes wrong, your company owns it. There is no putting that off onto a model provider, and I see nothing in any plausible future that changes it.

And the push factor here is not enthusiasm. It is competitive markets. Adoption cascades when your competitor’s unit economics start departing from yours, and by the time you can see that in their numbers they have already been through the expensive part of the learning.

The organizations spending now, even without clean ROI, are likely to be the ones that can show it over the next 24 months. Not because the spend was smart, but because they will have done the experimentation, made the organizational changes, and built the measurement discipline the answer requires.

The question is not whether AI pays. It is whether you will be able to prove it when someone asks.

Unlock your
true speed to scale

Accelerate what data and AI can do together.

Before you go - don’t miss what’s next in AI.

Stay ahead with Gruve’s monthly insights on trusted AI, enterprise data, and automation.