Enterprise AI Has Until January to Prove It Works
A year after MIT’s Project NANDA published the finding that 95% of enterprise generative-AI initiatives produced no measurable business return despite $30 Billion to $40 Billion of spending, the corporate record contains almost no trace of them. No cancellation notices, no write-downs, no post-mortems filed with investors. The study drew on 52 executive interviews, surveys of 153 leaders and analysis of 300 public deployments, and found that only 5% of integrated systems created significant value. The other 95% left no paper trail.
Capital has treated the finding as noise. Microsoft, Alphabet, Meta and Amazon have guided to a combined $725 Billion in capital expenditure for 2026, up 77% from roughly $410 Billion the previous year, and enterprise buyers have continued to sign. The gap between an industry that spends as though almost everything works and a research base showing almost nothing does rests on a bookkeeping convention that enterprise technology has never examined in public.
Initiatives that fail to deliver are rarely stopped. They are paused, re-scoped, absorbed into a wider transformation programme, reclassified retrospectively as a proof of concept, or carried into the following budget cycle under a different name and a different owner. The write-off column stays empty because nothing is ever formally written off, and the record of what went wrong stays with whoever happened to be in the room.
What makes the disappearance so complete is that the initiatives usually worked once. They performed in a pilot environment built to let them perform, with a curated dataset, a controlled scope and compliance questions chosen because they had answers. The failure came later, at the point where the system met the organisation as it actually operates, and by then the language for describing that failure had already been settled.
Nobody counts what nobody has defined
Joe Dunleavy, EMEA CTO and Head of Dava.X AI at Endava, declines to take the headline figure at face value, and his objection goes deeper than a defence of the industry. He wants to know what the 95% is counting.
“Seeing those numbers around the 95% and all that goes with it, I’m not quite sure the context of it is included,” he said. “Is it people building things using AI, or people using AI to build solutions?” The distinction has become material because enterprise licences have put model access into the hands of every function. A legal team builds a contract-review assistant. Twenty employees, working independently and unaware of one another, build twenty versions of the same expenses tool. Nineteen get abandoned within a month, and no register anywhere records that they existed.
“You’ve got anybody inside an organisation or anybody inside a customer base who can create something,” Dunleavy said, describing controls designed for a small central technology group and now stretched across an entire workforce. The organisation cannot separate technical incompetence from poor scoping, because it never had visibility of either. His own diagnosis of the underlying problem has been consistent since the first wave of deployments: people are brilliantly solving the wrong problem using AI, and doing so with real skill.
Jessica Constantinidis, Innovation Officer EMEA at ServiceNow, sees the same fog from the customer side, where the confusion starts with terminology and ends on the invoice. “A lot of people don’t even know that distinction,” she said of the line between automation, generative AI and agentic systems. “They push in stuff that doesn’t necessarily need AI, and they’re burning tokens because the prompting is wrong, the loops are wrong, the context is wrong.”
The pilot succeeds because the pilot was never the test
The gap between a working demonstration and a working deployment is where most of the 95% actually lives, and Constantinidis is unequivocal about what sits in that gap. “Data, data, data is the issue,” she said.
Her description of the mechanism explains why so many initiatives clear one bar and fail the next. “In a pilot you control the data. You have limited data and you just map it in, and you’re 100% sure it’s secure, it’s governed, it’s even maybe isolated from the production network,” she said. Production asks a different question entirely, one about intellectual property, compliance and governance across an environment nobody fully inventoried. No chief information security officer signs that off on optimism. “The investment has not caught up with the prerequisites of the data security and governance.”
The funding pattern compounds it, because enterprises have been paying for the visible layer while the invisible one degrades. “We’re so focused on the AI that we’re forgetting the baseline. The money is burnt on the AI and they don’t have money for the baseline,” Constantinidis said. Her recommended sequence runs the other way round: automate the operational drudgery first, capture the cost released by doing so, then spend that money on the harder work.
Levent Ergin, Chief Climate, Sustainability and AI Strategist at Informatica, adds a foundation most organisations assume they already have. Business process maps get built by analysts, go stale within months, and describe the idealised version of a workflow while the real one runs on email attachments and spreadsheets. “Most organisations don’t even have a good grasp on how things work today in their organisations,” he said. Any subsequent claim of measured productivity improvement becomes a comparison against a baseline nobody verified.
Dunleavy has watched the same sequence end the same way. Teams build something on a poor dataset, get excited, fail to get the output they want, and quietly pivot to something else. Nothing is recorded, and the underlying data problem survives to defeat the next attempt.
The vocabulary of the next budget cycle absorbs everything
What happens next is the part almost nobody documents, and the reason is not mysterious. “No one is willing to say we are not succeeding with AI, because everyone is trying to look good to their investors, to the board, to the public,” Ergin said. “Whenever AI doesn’t return on expectations, they say that was another proof of concept, or it gets absorbed by another transformation programme, and then it gets divided up into pieces.”
That division is where the audit trail ends. Laxmi Nageswari, Chief AI Officer at Cloud Box Technologies, describes the same disappearance from the delivery side. “Typically, failed AI initiatives are treated as pilot projects. It saves organisations from tagging these projects as failed,” she said. “The project failures are distributed across multiple areas such as cloud transformation, analytics, innovation budgets and IT modernisation, making it difficult to ascertain the financial returns of such investments.”
Keith Nolan, Director of Technology Delivery at C86, argues that the distortion begins upstream, in what enterprises agree to count as an achievement in the first place. “The market is currently overstating enterprise AI success by placing too much emphasis on the number of pilots launched, rather than the number of initiatives that have reached production, remained supportable and delivered measurable business outcomes,” he said. Work that stalls, in his account, “often remains hidden within broader innovation budgets rather than being recognised as discrete failures.”
Andrew Hewitt, Vice President Strategic Technology at TeamViewer, points to how difficult attribution becomes once an initiative sits inside a wider programme, which he says describes most of them. An HR team adding an AI assistant for new-hire questions is usually doing so alongside a new HR system and a revised onboarding process, and no review can isolate which change moved the number. Initiatives that fall short “quietly get scaled back, paused, or folded into a different initiative before anyone has to formally call it a loss,” he said. Handled badly, the ending is quieter still: “the initiative just fades from view, budget moves elsewhere, and the lessons live only with the people who worked on it.”
Nageswari is candid that the concealment has a rationale beyond vanity. Organisations protect reputation, customer trust and investor confidence, and the sums involved are large enough that disclosure carries consequences of its own. She argues that the calculation has been made too many times in a row, and that the industry now prices a technology it cannot see clearly.
The finance function has started asking, and the reporting cannot answer
For most of the past two years the spending went unexamined, which Dunleavy recalls with something close to embarrassment. “We finally stopped seeing a celebration of tokens spent,” he said, describing the internal consumption leaderboards of a year ago as an obvious waste of money. Chief financial officers have replaced the leaderboards with questions, and the questions have exposed how little instrumentation exists.
“You can only measure it if you can actually get an accurate picture, and I don’t believe we have an accurate picture,” Dunleavy said. Enterprise licensing products lack the reporting sophistication that buyers take for granted in other categories, which leaves organisations pressing providers for usage data, spend visibility and return projections they cannot currently obtain.
Constantinidis reports the same blind spot inside customer organisations, where nobody can attribute cost to a department, and where agents commissioned months ago keep running in the background consuming tokens after everyone involved in creating them has moved on. Her prescription is to onboard and offboard AI agents with the same discipline applied to employees, and to run dashboards capable of surfacing duplicates before they multiply. She draws the parallel with identity and access management, a problem the industry took two decades to take seriously.
Jennifer Dsilva, VP Finance at Omnix International, treats the problem as a governance failure with a finance signature. “While undertaking an AI initiative, it is important to get the budget approved in advance, measure the expected return on investment to consider this as a capital expenditure or expense,” she said. Her verdict on causation leaves the technology largely out of it: “AI initiatives usually fail less because the technology is weak, and more because adoption, training, use-case definition, ownership, and measurement are weak.”
Success criteria written afterwards are not criteria
Every leader interviewed for this piece arrived at the same corrective from a different direction. Dunleavy rejects organisation-wide productivity mandates outright, because productivity means something different for a legal team than for a customer service desk. “Pick the two to three areas of your organisation you want to improve productivity, and then what does that improved productivity look like for them? It needs to be OKR and KPI led,” he said. Where a company wants a broad rollout anyway, simply to get people using the tools, he advises capping the spend and funding a real training programme, “because otherwise you’re just burning your money.”
Constantinidis has watched the alternative play out in meeting after meeting. “I said to them, how are you using AI? They said, we’re doing this and this and this. And I said, so what is the business goal? They are in headlights,” she said. The answer she wants is narrow by design: cost reduction, cost avoidance, faster go-to-market, or margin uplift. Her comparison with conventional delivery is unflattering. In any standard project there is a budgeting exercise, and for AI work there is none.
Dsilva offers what a defensible target looks like in practice, citing a reporting cycle cut from two days to two hours, an outcome that can be verified rather than asserted. Nolan sets out the staged version of the same discipline, arguing that organisations should validate the commercial hypothesis, prepare the operating model, deploy into production, demonstrate measurable outcomes, then prove the capability scales securely and economically. “A successful proof of concept is not the same as a successful investment,” he said.
Ergin puts ownership at the centre of it, noting that companies are increasingly appointing a chief AI officer with structural responsibility for a portfolio rather than a single initiative, precisely so that someone is accountable for reviewing what failed. He also draws a line between external and internal candour. Control the narrative for investors and the public, he said, while keeping the internal account honest: “The narrative internally should be very clear on what went wrong and why, so that you can learn from it.”
Mena Migally, Regional Vice President for EMEA East at Veeam, situates the current moment inside the pattern of earlier technology cycles, and treats pilots that miss their targets as inputs into better data governance and operational readiness rather than sunk cost. He argues for defining success metrics before deployment across revenue growth, operational efficiency, risk reduction and recovery capability, and for telling employees plainly what an initiative is intended to achieve, on the basis that distrust of AI inside a workforce is usually a response to opacity rather than to the technology.
Dunleavy extends that point into territory most organisations avoid. Where AI is being deployed to reduce headcount, he argues, the objective should be stated openly and tracked as a formal target, because employees who suspect a stealth programme will resist it exactly as they would resist any other change imposed without explanation.
January 2027 is when the accounting starts
The wider evidence suggests the silence has a shelf life. S&P Global Market Intelligence found that 42% of companies abandoned most of their AI initiatives in 2025, up sharply from 17% a year earlier, while RAND Corporation research puts the AI project failure rate above 80%, roughly twice that of conventional IT projects. Gartner expects more than 40% of agentic AI projects to be cancelled by the end of 2027 on escalating costs, unclear business value or inadequate risk controls, and reckons only around 130 of the thousands of vendors claiming agentic capability are genuine.
Hewitt’s research explains why the internal ledger looks so strange. Just 3% of employees report no noticeable benefit from AI at work, while 56% verify AI output often or always, spending close to two hours a week doing it. Value that is real and diffuse never reaches the profit and loss account in a form a review will accept, while the labour cost of governing it goes uncounted on the other side. His research also found 71% of managers describing their experience with AI as positive against 58% of employees, a gap that suggests the people reporting the returns are not the people generating them.
Nageswari argues the concealment carries a cost the whole market pays. Publishing both outcomes, she said, “helps set up realistic goals and expectations related to cost of implementation, deployment timelines, and returns on investment.” Nolan puts the equivalent case to vendors: “Reporting dozens of pilots tells the market very little; reporting how many reached production, remained operational and delivered measurable value is far more useful.”
Dsilva makes the case from inside the finance function, where the appetite for candour has a practical limit and a practical purpose. AI is a disruption, she said, and organisations need to give their people the flexibility to fail, learn, adapt and retry, provided the failures are accounted for operationally as well as financially.
Dunleavy expects the change to come through procurement rather than conscience. “If all you do is share across your organisation what’s been successful, then you’re not learning from the stuff that hasn’t been successful,” he said, describing the alternative as the definition of insanity, running the same approach against the same broken data and expecting a different result. His forecast for the coming months is a race towards compliance, regulation and reporting insight that has never existed, driven by return-on-investment scrutiny and a tightening regulatory environment in equal measure, with the EU AI Act removing the defence that an application was built by somebody outside the technology organisation.
Ergin’s closing argument is that the winners will be identifiable well before the numbers confirm it. Companies that obsess over controls, governance and business understanding, rather than chasing the latest marketing trend, will be the ones to convert experimentation into value.
The deadline is commercial. When enterprise contract renewals come round in January 2027, buyers will want evidence the market does not currently produce. At that point the 95% stops being a statistic about last year’s pilots and becomes a question the finance function asks before signing anything.