Whose Failure Is the AI ROI Miss?
Only 29% of enterprises report real ROI from generative AI, and just 23% say the same about AI agents (The State of AI Adoption in the Enterprise [Q1 2026 Review]). Gartner expects more than 40% of agentic AI projects to be canceled by 2027 (Why 40% Of Agentic AI Projects May Be Canceled By 2027).
Numbers like that used to trouble me. Not so much anymore. What they mostly do now is make me curious about a smaller, meaner question: who ends up taking the blame? That is usually how downturns finish, with someone becoming the reason, even when the real reason is everyone.
I present an argument in this article from four different angles:
- Finance
- Product
- Engineering
- Governance.
Curious? Read on ...
The numbers nobody is disputing
Roughly $2.52 trillion is heading into AI in 2026, a jump of 44% year over year, per Gartner: Global AI spending to reach $2.5 trillion in 2026, and yet 79% of them still say they're running into real adoption challenges, according to the Writer Survey: 79% of Enterprises Face AI Adoption Challenges Despite Record Investment. Gartner points to three reasons behind the cancellations: cost that keeps climbing, business value nobody can quite pin down, and risk controls that simply aren't where they need to be.
Model selection/quality, tools, harnesses matter. But they are not the primary, hard problems that are causing these ROI issues.
The first two functions: The money people and the product people, and they are both right!
The finance argument is the cleanest one to make. Nobody built a credible ROI model before the budget was approved. That is the whole story in one line. As CIO.com writes: "AI isn't failing; companies are". That claim aligns well with a pattern I am noticing. Workloads carrying a pre-defined success metric and an existing audit trail (e.g. fraud detection, invoice matching, ticket triage) tends to keep their funding; open-ended "AI transformation" programs mostly does not (State of the CIO, 2026: CIOs set the course for AI ROI).
Coding and engineering alone pull roughly 55% of departmental agentic spend (55% of All Departmental AI Spend Is Now on Coding). Not because coding tools carry more magic than the rest of the portfolio. It is just easier to measure - a number a CFO can drop straight into a spreadsheet.
The product argument, oddly enough, leans on the exact same evidence. But the terminology used by Product teams and Finance teams are different. One blames the lack of ROI, the other blames the lack of value and outcome. But ultimately they are two sides of the same coin.
Gartner's own analyst calls most agentic deployments "early-stage experiments... driven by hype and often misapplied," and reckons only about 130 of the thousands of vendors marketing "agentic AI" have any real capability behind the label. The rest is what the industry now calls agent-washing, and that is not a compliment. CIO.com went further, finding that most enterprise AI investment runs with no executive owner at all.
You cannot fail a test you never wrote! Similarly, a program cancelled for "unclear business value" is difficult to perform a RCA on considering it did not have a real value analysis done in the first place.
The third function: Engineering. The failure no funding gate ever catches!
Then comes the engineering argument. Researchers have catalogued 63 confirmed budget-overrun incidents spread across 21 orchestration frameworks between 2023 and 2026 (Token Budgets: An Empirical Catalog of 63 LLM-Agent Budget-Overrun Incidents): agents that just keep retrying a broken tool call, over and over, without ever stopping. No pre-funding ROI model was ever going to catch a loop that starts nine months into production. That single concession, I think, cracked open the most interesting question in the whole debate.
Is a spend ceiling a governance requirement, or is it an engineering build?
My answer is both, and the two halves do not share an owner. Mandating that a ceiling has to exist is cheap. It is real, and it is enforceable right at the funding gate. Building a ceiling that survives concurrent agents, that does not lose a race condition under load, and that cannot be reasoned around by the very model it is meant to restrain is systems engineering. Full stop. One documented team cut its monthly agent spend from $87,000 down to $24,000 once real controls actually went in (AI Agents Burn 50x More Tokens Than Chats), and the gap between having that layer and not having it is nowhere close to a rounding error. This argument, in the end, shows up on an invoice.
And the fourth funciton: Governance. They can mandate the box, but cannot fill it.

That same split turned up again, but this time one level higher.
Once a mandated, centrally engineered enforcement layer had settled the spend-ceiling argument, it grew tempting to fold product validation into governance too, requiring a value-hypothesis document inside the same review that certifies the architecture. I like the instinct behind it; it fails for the same reason the checklist-based spend gate failed before it, and for a familiar reason: the box was never the problem.
A review board can verify that a document exists. What it cannot confirm is whether the customer described inside is real, or whether the chosen metric is even the right one.
A team whose funding depends on chasing the word "agent" will write a believable hypothesis just to clear the gate. So, that won't work in many cases.
What a board can actually do, portfolio-wide, is mandate that every project names an accountable "Value and ROI Owner" before funding is released. Since no governance committee member's bonus ever rides on whether the metric actually moves, they should not be the ones taking accountability.
Four jobs, not one villain
Strip away who is arguing which position, and the four cases start converging on something far more useful than a single guilty party. Four jobs. Not one villain.
Capital allocation, product definition, engineering operability, and architectural governance are four distinct jobs, and none of them substitutes for the others.
Once you look at it that way, the ROI numbers make a different kind of sense. All four jobs went unfilled at once: a funding gate with no kill criterion, a roadmap with no validated job sitting behind it, nobody wired a circuit breaker into the production system, and an estate with no shared definition of what "agentic" is even allowed to mean. Four gaps, not one.
The workloads that are still getting funded tell the same story from the other direction. Fraud detection. Customer service deflection. Supply chain optimization. Developer tooling, kept narrow in scope. These are not the lucky exceptions; they are simply the cases where enough of those four jobs happened to get done at once.
What I find more useful than picking one scapegoat to rotate through every quarter is this: fix the job that is actually missing, and the numbers follow.
What I would actually ask for
If I were running an AI portfolio right now, I would stop letting "we have a policy for that" pass as a finished sentence. Not on its own. Each of the four jobs above needs a paired commitment, and that pairing has two very different halves: a mandate on one side, and on the other, someone who actually built the thing, tested it, and now owns the enforcement. Gartner's 2027 cancellation wave will be a rough natural experiment, the kind that shows who actually did the work. The organizations that closed that gap will look very different from the ones that only ever got as far as the policy document.
I would watch closely which failure mode dominates the postmortems: cost, value, or risk controls.
That detail will say a lot about which job actually got fixed first, and just as much about which one most people quietly assumed someone else was handling.
The question worth asking instead

"Whose failure is it" is the wrong question to walk away from this with. Wrong question, full stop.
The better one is narrower, and it is more useful to ask on a Monday morning: for each control this technology requires, a budget, a validated job, a safe runtime, a shared standard, does it merely exist on paper, or does it actually hold up under load, with someone accountable for the gap between the two?
I believe the enterprises that can answer that question honestly, job by job, not just once but continuously, will be the ones still running their agentic projects in 2028; the rest will make up Gartner's 40%, filing a postmortem.
Citations & Further Reading
- The State of AI Adoption in the Enterprise [Q1 2026 Review]
- Why 40% Of Agentic AI Projects May Be Canceled By 2027 (Forbes)
- Gartner: Global AI spending to reach $2.5 trillion in 2026 (Computerworld)
- Writer Survey: 79% of Enterprises Face AI Adoption Challenges Despite Record Investment (Welcome.AI)
- Why enterprises aren't seeing AI ROI, and what CIOs can do about it (CIO.com)
- State of the CIO, 2026: CIOs set the course for AI ROI (CIO.com)
- 55% of All Departmental AI Spend Is Now on Coding. And It's Not Slowing Down (SaaStr)
- Gartner: 40% of agentic AI projects will fail, making humans indispensable (MarTech)
- Token Budgets: An Empirical Catalog of 63 LLM-Agent Budget-Overrun Incidents, with an Affine-Typed Rust Mitigation as a Case Study (arXiv)
- AI Agents Burn 50x More Tokens Than Chats (LeanOps)
Comments ()