Skip to content

AI ROI: What the Published Evidence Actually Shows

Reported AI ROI ranges from 3.7 dollars per dollar to a 95% failure rate. The gap is definitional, not factual. A reconciliation of the primary studies, what each one counted, and how to measure return on your own AI work without fooling yourself.

Occasional field notes on building software, no spam

Protected by Cloudflare Turnstile · Privacy · Terms

Idealogic: AI ROI

Two numbers dominate almost every conversation about AI ROI, and they point in opposite directions. An IDC study sponsored by Microsoft put the average return at 3.70 dollars for every dollar invested in generative AI. A report out of MIT the following summer reported that 95% of enterprise generative AI pilots produced no measurable bottom-line impact. Both get quoted as though they settle the question. Neither does, because they are not measuring the same thing, on the same population, with the same definition of success.

This piece assembles the published evidence on AI return in one place, states plainly where it disagrees with itself, and works out which disagreements are real and which are artifacts of how the question was asked. Then it gets practical: what the real cost of an AI project is, where AI ROI genuinely comes from, and how to measure it on your own work without fooling yourself. Every figure below carries its publisher, its date, and its sample, because in this topic the sample is usually the whole story.

The short version

  • Reported AI failure rates run from 20% to 95%. The spread is almost entirely definitional: what got counted (a pilot, a use case, a company), what counted as failure, and who was asked.
  • The two most-quoted opposite numbers are closer than they look. McKinsey's 2025 survey found about 6% of organizations reach a 5% or greater EBIT impact from AI. MIT NANDA reported about 5% of pilots producing real value. Held to the same threshold, a global executive survey and a disputed deployment review land within a percentage point of each other.
  • Almost every headline AI ROI figure is self-reported, and self-report is systematically unreliable here. Wharton's ROI question asks executives what they have heard "based on internal conversations." BCG states in its own methodology that its numbers "may be subject to perception bias."
  • Where return has been measured rather than asked about, the results split by population. Customer-support agents got 15% faster. Experienced developers working in codebases they knew got 19% slower.
  • The cost side is dominated by data work, integration, evaluation, governance, change management, and recurring inference. The model is usually the smallest line item.
  • The measurement discipline that makes AI ROI provable, a baseline captured before the build and a real counterfactual, costs almost nothing and is skipped almost universally.

What the published record on AI ROI actually reports

The published record on AI ROI is a set of surveys that disagree with each other by a factor of four, plus a much smaller set of measured studies that disagree by sign. Here is what each one actually found, with what it counted.

The two AI ROI figures everyone quotes

IDC, sponsored by Microsoft, November 2024. The InfoBrief "2024 Business Opportunity of AI" (IDC# US52699124) surveyed more than 4,000 business leaders and AI decision-makers worldwide. Its headline: "For every $1 a company invests in generative AI, the ROI is $3.7x," with top leaders at 10.3, value realized within 13 months, and deployments taking under eight months. The public summary does not disclose fieldwork dates, country breakdown, or how ROI was computed by respondents. It is a vendor-sponsored study, and the sponsor sells the product.

MIT Project NANDA, July 2025. "The GenAI Divide: State of AI in Business 2025," by Aditya Challapally, Chris Pease, Ramesh Raskar, and Pradyumna Chari, produced the 95% figure that has done more work in boardrooms than any other number in this field. It deserves a caveat that almost nobody attaches to it. The report's official address, nanda.media.mit.edu/ai_report_2025.pdf, now returns an HTTP 302 redirect to the NANDA project overview page, which does not host the report (checked 4 August 2026). Contemporaneous accounts of its methodology do not agree: some describe 52 executive interviews, 153 survey responses, and 300-plus implementation reviews, while Fortune's report describes 150 interviews with leaders, a survey of 350 employees, and an analysis of 300 public deployments. The work was never peer reviewed. The finding may well be directionally right, and the section below argues it is, but it should be cited as a preliminary industry report with a withdrawn primary document, not as an MIT study in the sense that phrase usually implies.

The five surveys that fill in the middle of the AI ROI range

S&P Global Market Intelligence, May 2025. "Voice of the Enterprise: AI & Machine Learning, Use Cases 2025" surveyed 1,006 midlevel and senior IT and line-of-business professionals across North America and Europe. It found that the share of companies abandoning most of their AI initiatives before production rose from 17% to 42% year over year, with organizations reporting on average that 46% of projects are scrapped between proof of concept and broad adoption.

IBM Institute for Business Value, May 2025. The 2025 CEO Study, run with Oxford Economics across 2,000 CEOs in 33 countries and 24 industries, found only 25% of AI initiatives delivered expected ROI over the last few years, and only 16% scaled enterprise-wide. The same CEOs expect 85% of their scaled efficiency investments to be ROI-positive by 2027.

BCG, September 2025. "The Widening AI Value Gap" reports on the Build for the Future 2025 Global Study: 1,250 CxOs and senior executives who are AI decision makers, drawn from 68 countries and more than 25 sectors. Only 5% of firms are "future-built" and achieving AI value at scale, while 60% are "reaping hardly any material value, reporting minimal revenue and cost gains despite substantial investment".

McKinsey, November 2025. The State of AI survey was in the field from 25 June to 29 July 2025 and drew 1,993 responses from 105 nations, weighted by each nation's contribution to global GDP. It found 88% of respondents reporting regular AI use in at least one business function, and 39% attributing any level of EBIT impact to AI at the enterprise level. Most of that 39% put the impact under 5%. About 6% of respondents qualify as high performers, meaning 5% or more of EBIT attributable to AI plus self-reported significant value.

Gartner, April 2026. A survey of 782 infrastructure and operations leaders, conducted in November and December 2025, found that only 28% of AI use cases in I&O fully succeed and meet ROI expectations, while 20% fail outright. The number that gets quoted from this release is the 28%. The number nobody quotes is the 52% in between, which is the modal outcome: partially working, not clearly failed, not clearly paying.

Bar chart comparing reported AI ROI failure rates across six studies, from Gartner's 20% outright failure rate for infrastructure and operations use cases to MIT NANDA's 95% of pilots with no measurable P&L impact.
Same topic, six studies, a 75-point spread. Each bar counts something different
Source and dateReported figureWhat was countedFailure defined asSample
Gartner, Apr 202620% fail, 28% fully succeedAI use case in I&Ooutright failure782 I&O leaders
S&P Global, May 202542%companyabandoned most initiatives pre-production1,006 IT and business professionals
BCG, Sep 202560%companyno material revenue or cost gain1,250 CxOs, 68 countries
McKinsey, Nov 202561%companyno enterprise-level EBIT impact1,993 executives, 105 nations
IBM IBV, May 202575%AI initiativedid not deliver expected ROI2,000 CEOs, 33 countries
MIT NANDA, Jul 202595%pilotno measurable P&L impactdisputed, primary withdrawn

Read down that table and the apparent chaos organizes itself. Nobody is lying. The counters counted different objects.

How the 95% failure rate and the 3.7x return fit together

The two figures describe different denominators and different thresholds, and once you hold both constant, the studies converge to a degree that ought to be more famous than the argument.

Start with the threshold. MIT NANDA counted pilots that produced no measurable P&L impact, and put the survivors at about 5%. McKinsey asked executives about EBIT attributable to AI at the enterprise level, and found about 6% clearing a 5%-of-EBIT bar with self-reported significant value. Those are close to the same question: how many organizations can point at the income statement and show AI on it. One study is a preliminary review of roughly 300 deployments with a contested method. The other is a 1,993-respondent survey across 105 nations, run annually for years by a firm with no incentive to depress the number. They agree within a point.

Now the denominator. The 3.7x is not a failure rate at all. It is an average return reported by the subset of people who had a return to report, in a study commissioned by a company selling AI products. Averages over self-selected successes are not comparable to base rates over all attempts. If 5% of efforts produce a strong return and 95% produce nothing, and you ask the population what return they saw, the answers cluster among the people with an answer. Both numbers can be exactly right and describe the same world.

The S&P Global survey makes this visible inside a single dataset. The same 1,006 respondents who reported abandoning 42% of their initiatives before production also reported positive impact on revenue growth objectives at 76%, cost management at 74%, and risk management at 70%. Those figures are not in tension. A company can scrap most of what it starts and still say the surviving work helped. The failure rate and the satisfaction rate are answers to different questions, and the same firm can honestly give both.

The 95% and the 3.7x are not a contradiction to resolve. They are a base rate and a conditional average, quoted as though they were the same statistic.

There is a second reconciliation hiding in the S&P data, and it is less comfortable. In the same survey, 46% of respondents whose organizations had invested in generative AI reported that no single enterprise objective had seen a strong positive impact. Roughly half say it helped somewhat everywhere and strongly nowhere. That is the honest center of the distribution, and it is the outcome no headline describes.

Why measured AI return varies so much between studies

Measured AI numbers vary because four things vary underneath them: the unit of analysis, the question wording, who is in a position to know the answer, and what the respondent is rewarded for saying. This is not speculation. The Federal Reserve published a note working through exactly this problem for AI adoption, and its framework transfers directly to ROI.

In "Monitoring AI Adoption in the U.S. Economy" (FEDS Notes, 3 April 2026), Jeffrey S. Allen compares three instruments measuring the same phenomenon. The Census Bureau's Business Trends and Outlook Survey put firm-level AI adoption at about 18% at the end of 2025. The Real-Time Population Survey put work-related generative AI use at about 41% of individuals in November 2025. The Survey of Business Uncertainty, run by the Atlanta Fed with Chicago Booth and Stanford, found 78% of the labor force works at firms that have adopted AI. Three credible instruments, a four-fold spread, no error anywhere.

Bar chart showing four measures of AI adoption used as denominators in AI ROI research: 18 percent of US firms in the Census BTOS, 41 percent of workers in the Real-Time Population Survey, 78 percent of the labor force at AI-adopting firms in the Survey of Business Uncertainty, and 88 percent of organizations in McKinsey's executive survey.
Four instruments, one phenomenon, an 18 to 88 percent range. Sampling frame explains most of it

Four reasons AI ROI numbers disagree

Allen's four mechanisms are worth reading against any AI ROI figure you encounter.

The sampling frame decides the answer. The Census survey draws on roughly 1.2 million businesses split across six panels, which means it reflects the actual population of US firms, and most US firms are small. The Survey of Business Uncertainty deliberately oversamples large employers and runs on about 150 responses a month. McKinsey's panel skews large too: 38% of its respondents work at organizations above 1 billion dollars in revenue. Ask large firms and you get large-firm answers. As of the collection period ending 3 May 2026, the Census Bureau put overall AI use at 19.8%, against about 37% for businesses with at least 250 employees and under 20% for those with four or fewer.

Question wording moves the number, sometimes more than reality does. The Census Bureau originally asked about AI use "in producing goods or services," a deliberately conservative framing. In November 2025 it broadened the question to use "in any of its business functions," and Allen points to that revision as a likely explanation for the jump in reported adoption that followed. Any AI ROI survey that asks "has AI delivered value" will produce a higher number than one that asks "what percentage of EBIT is attributable to AI," and the difference tells you about the questionnaire.

The respondent may not know. Allen notes that around 10% to 11% of Census respondents answer "do not know" on AI questions, and argues that business executives in the Survey of Business Uncertainty are better placed to answer for the firm. The same asymmetry runs through the AI ROI literature in the other direction: the executive who authorized the AI budget is well placed to describe the strategy and poorly placed to know whether the support queue actually got faster.

Social desirability pressure has flipped direction. Allen makes a point that is easy to miss and hard to unsee: early in the generative AI cycle, firms were cautious about admitting AI use because of security concerns, while more recently "firm representatives responding to surveys may face pressure to report AI usage as an efficiency initiative." The incentive to under-report became an incentive to over-report inside about two years. Time series across that window are measuring the incentive as well as the behavior.

The self-report problem in AI ROI surveys

Almost every widely cited AI ROI number is a self-report, and two of the studies say so themselves.

Wharton Human-AI Research and GBK Collective's "Accountable Acceleration" (October 2025) is a 15-minute online survey of roughly 800 senior decision makers at US enterprises with over 1,000 employees, fielded between 26 June and 11 July 2025. It reports that 74% see positive ROI, a figure widely repeated as evidence the return is real. The question that produced it reads: "Based on internal conversations with colleagues and senior leadership, what has been the return on investment (ROI) from your organization's Gen AI initiatives to date?" That is a measure of what executives have heard, and the report is admirably transparent about it.

The seniority gradient in the same data makes the point sharper. Among VP and above, 45% called ROI significantly positive. Among managers and directors, 27% did. Mid-managers were twice as likely to answer "too early to tell." These are people inside the same firms describing the same initiatives, and their answers differ by 18 percentage points along the org chart. Whatever that gradient measures, it is not the P&L.

BCG puts the caveat in its own methodology section. Its value figures come "from self-reported data provided by the respondents," the results reflect the business area each respondent knows best rather than the whole company, and "responses may be subject to perception bias." That is a research firm telling you, in the appendix, how to read its own headline.

What happens when you measure AI return instead of asking about it

When researchers measure output directly instead of asking people how they feel about it, the numbers get smaller, more specific, and occasionally negative. Three studies define the current state of that evidence, and they do not agree, which is what makes them useful.

Customer support, positive and large. Erik Brynjolfsson, Danielle Li, and Lindsey Raymond studied the staggered rollout of a generative AI assistant across 5,172 customer-support agents, published as "Generative AI at Work" in the Quarterly Journal of Economics (vol. 140, no. 2, 2025, pp. 889 to 942). Access to AI raised issues resolved per hour by 15% on average. The distribution matters more than the average: less experienced and lower-skilled workers improved both speed and quality, while the most experienced and highest-skilled agents saw small gains in speed and small declines in quality. The tool compresses the gap between novices and experts by lifting the bottom, not the top.

Experienced software developers, negative. METR ran a randomized controlled trial with 16 experienced open-source developers completing 246 tasks in mature repositories where they averaged five years of prior experience, published as arXiv:2507.09089 in July 2025. Before starting, the developers forecast that AI would cut completion time by 24%. Afterwards, they estimated it had cut completion time by 20%. Measured, allowing AI increased completion time by 19%. Economics experts had predicted a 39% speedup and machine learning experts 38%. METR is careful that its developers and repositories do not represent a majority of software work, and the organization has since revised its experimental design, so treat the result as a well-executed finding about a narrow population rather than a verdict on AI coding tools generally.

Set the two side by side and the mechanism is legible. The support study found the largest gains among people with the least context. The developer study found losses among people with the most. Both are consistent with a tool that supplies competent generic knowledge quickly: enormously useful when your baseline knowledge is thin, a tax on your time when you already hold five years of a codebase in your head. Any AI ROI estimate that ignores who is using the tool is estimating the wrong thing. Stanford HAI's 2026 AI Index summarizes the pattern across the literature as gains of 14% to 15% in customer support, 26% in software development, and 50% in marketing output, with smaller gains in tasks that require deeper reasoning.

The macro view, effectively zero so far. Anders Humlum and Emilie Vestergaard linked two large adoption surveys, each covering roughly 25,000 workers at 7,000 workplaces across 11 AI-exposed occupations, to Danish administrative records. Their July 2025 paper, then titled "Large Language Models, Small Labor Market Effects" and circulated as NBER working paper 33777, estimates "precise null effects at both the individual and workplace levels, with confidence intervals ruling out effects above 1%" on earnings and recorded hours, alongside productivity gains they put at 3% in time savings. The paper has since been revised and retitled, so check which version you are citing.

That last study is the one worth sitting with, because it is the closest thing we have to a controlled read on whether micro-level time savings become money. Over two years, in a country with excellent administrative data, they did not. The authors attribute this partly to new integration and oversight tasks absorbing the saved time, and frame the result as the slow front end of a productivity J-curve rather than a permanent null. Either way, the inference for a business case is direct: hours saved are a leading indicator, not a realized return, and the conversion is neither automatic nor fast.

The real cost of an AI project, beyond the model

The cost of an AI project is dominated by everything that is not the model. This is the denominator problem, and it invalidates more AI ROI claims than any error in the numerator.

The model is the cheap part. A hosted API or a licence is a known, often modest line. It feels like the project because it is what produces the magic in the demo. In production it is frequently the smallest cost of the lot.

Data work is usually the biggest hidden cost. An AI system is only as good as what it can read, and most enterprise data is scattered, stale, inconsistent, or locked in systems nobody can reach cleanly. Gartner's April 2026 I&O survey found 38% of leaders who hit setbacks naming poor data quality or limited data availability as a direct cause of failure, tied with skills gaps at 38%.

Integration is ordinary production engineering. A model in a notebook touches nothing and returns nothing. Wiring it into the CRM, the support desk, the claims pipeline, or the codebase means APIs, authentication, error handling, and the unglamorous decisions that connect a model to a live workflow. It costs what production engineering costs. Gartner's I&O data supports the point from the other side: among leaders who delivered at least one successful AI use case, success was attributed primarily to integrating AI into existing workflows and systems.

Evaluation is not optional and not free. You cannot prove a return on output nobody verified. A test set of real cases, defined quality criteria, and a way to catch regressions is a real build. Skipping it is how teams ship on impressions and learn about problems from angry users.

Governance carries its own cost. The moment an AI system touches real data or takes real actions, it needs access controls, permission scoping, approval steps, and logging. McKinsey found 51% of respondents from organizations using AI reported at least one negative consequence, with roughly a third citing consequences from AI inaccuracy. Those consequences have a cost, and it belongs in the denominator.

Change management is the cost people forget entirely. A tool nobody adopts returns nothing, however good the model is. Training, clear ownership, honest communication about what the system can and cannot do, and a feedback loop are real work.

Inference is the recurring bill that erodes the return after launch. A normal feature is built once and runs nearly free. An AI feature charges you for every prediction, forever. When Gartner predicted in July 2024 that at least 30% of generative AI projects would be abandoned after proof of concept by the end of 2025, it named escalating costs alongside poor data quality, inadequate risk controls, and unclear business value. A thin per-transaction margin gets eaten by a per-transaction cost that never goes away.

Cost componentWhat it coversVisible at proposal?
The modelAPI calls, licences, per-token priceYes, and over-weighted
DataSourcing, cleaning, structuring, permissionsRarely, often the largest hidden cost
IntegrationWiring AI into live systems and workflowsPartly, underestimated
EvaluationMeasuring output quality at scaleAlmost never
GovernanceAccess, approval, logging, complianceAlmost never, until an audit
Change managementTraining, ownership, adoptionAlmost never
InferenceRecurring cost of every prediction in productionAlmost never, the post-launch surprise

If you take one thing from this section, take the denominator. A pitch built on the cost of the model is quoting a fraction of the real number. We unpack the full picture in our companion guide to what AI development actually costs, because getting the cost side right is half of getting AI ROI right.

Where AI ROI actually comes from

AI ROI comes from four sources, and being clear about which one you are chasing is the difference between a measurable result and a vague hope. They are listed in rough order of how reliably teams realize them.

Cost and time savings on high-volume repetitive work. The most dependable source and the easiest to measure. Support triage, document processing, first-draft generation, code assistance: work that is high-volume, somewhat repetitive, and tolerant of a verification step. McKinsey's 2025 survey found cost benefits from individual use cases concentrated in software engineering, manufacturing, and IT. The Brynjolfsson result sits squarely in this category, and so does its warning: the gain concentrates among your less experienced people.

Revenue lift in the right functions. McKinsey found revenue increases most commonly reported in marketing and sales, strategy and corporate finance, and product and service development, consistent across years of the same survey. Revenue attribution is harder than cost attribution, which is why a clean comparison matters more here than anywhere.

Risk and error reduction. Fraud detection, compliance monitoring, quality control, anomaly spotting: places where catching the bad case earlier has a clear financial value. The return is avoided loss, which is genuine but requires you to have measured the loss rate before, or it looks invisible. This is often the highest-value source in regulated industries and the one most likely to go uncounted, because nothing visible happened.

New capability. Occasionally AI lets a team do work it could not do before, at a volume or speed that opens something new. This is the hardest to put a number on and the easiest to over-claim, so treat it as upside rather than as the basis of the business case.

Source of returnTypical homeHow measurable
Cost and time savingsSupport, ops, document work, code assistanceHigh, baseline usually known
Revenue liftMarketing, sales, productMedium, attribution is the work
Risk and error reductionFraud, compliance, qualityMedium, needs a prior loss rate
New capabilityWhatever was previously impossibleLow, treat as upside

One pattern runs through all four, and it is the most consistently replicated finding in the whole literature. McKinsey reports that AI high performers are nearly three times as likely as others to say their organizations have fundamentally redesigned individual workflows, and that this redesign "has one of the strongest contributions to achieving meaningful business impact of all the factors tested." BCG's separate study of 1,250 executives reaches the same conclusion from a different direction, finding that its top 5% are distinguished by ambitious top-down programs with cost and revenue targets set by senior management, not by better models. The model is available to everyone at the same price. The redesign is what almost nobody does, and it is the closest thing to a causal story the survey evidence supports.

How to measure AI ROI without fooling yourself

Measuring AI ROI honestly requires four things: a baseline captured before the build, a defined attribution window, a counterfactual, and a fully loaded cost. None of it is exotic. All of it has to happen before you start, which is exactly when it feels least urgent.

1. Pick one use case with a measurable outcome. Resist the urge to measure the return on AI in general, which is unanswerable. Choose a single, bounded job where success has a number attached, defined before anything is built. "Cut average handle time on support tickets" is measurable. "Add AI" is not. The sharpness of the outcome is the precondition for everything else, and it is the same discipline that separates an AI implementation that reaches production from one that stalls in pilot.

2. Capture the baseline before you build. Measure the current state of your chosen metric while the AI is still hypothetical, and measure it long enough to see its normal variation. A single pre-period number is not a baseline, it is one sample from a noisy distribution, and support volumes have weekly and seasonal patterns big enough to swamp a real effect in either direction. Twelve weeks of history is a reasonable floor for anything with a weekly cycle. If you cannot get history, you can still capture a parallel period, which is worse but honest.

3. Define the attribution window before you look at the results. Decide up front how long after launch you will measure, and commit to it. Deciding afterwards is how a project that looked flat at 30 days gets re-measured at 90 until it looks good. The window should be long enough for adoption to stabilize, because early numbers measure novelty, and short enough that the rest of the business has not changed underneath you. For most workflow-level use cases that lands somewhere between eight and sixteen weeks post-adoption, with adoption itself measured, not assumed.

4. Build the counterfactual into the rollout. Attribution is the hard part of AI ROI, and the only clean answer is a comparison group. Three designs work in ordinary corporate conditions, in descending order of strength. Randomized assignment at the level of the ticket, the document, or the task is best, and it is what makes the METR trial hard to argue with: each task was randomly assigned to allow or disallow AI. A staggered rollout is next: release the tool to teams in an order that has nothing to do with how well they are performing, so earlier teams act as treatment and later ones as controls. That is the structure Brynjolfsson and colleagues exploited, and it is usually available for free, because most companies roll out in waves anyway. A matched holdout is third: keep a comparable team, region, or shift on the old process for the duration of the window.

5. Count the fully loaded cost, over a defined period. Add every line from the cost section, including the recurring inference bill, and express it as total cost of ownership over the same period as the benefit, usually a year. Inference makes AI an ongoing expense rather than a one-time build, so a first-year ROI and a steady-state ROI are different numbers and both are worth computing.

6. Track leading indicators alongside the financial result. Cycle time, error rate, output volume, and adoption move first and tell you early whether the thing is working. They are also the honest place to report time savings, which the Danish evidence suggests should not be booked as money until you can show where the money landed.

Need an AI use case that can prove its return?
We scope AI work the way it has to be measured: one sharp use case, a baseline captured before the build, a rollout designed so the comparison is real, and the integration and evaluation that make the return provable. Production is the bar, and a defensible number is the proof.
See how we ship AI

A measured 1.6x on a use case you instrumented properly is worth more to a board than a claimed 10x nobody can reconstruct, because the first one you can repeat. Where a single use case sits in a larger plan is the subject of building an enterprise AI strategy, and the groundwork that makes any of this possible is covered in our guide to AI readiness.

Measuring AI ROI when you cannot run a control group

Often you cannot hold anything back, and the AI ROI question does not go away because of it. The tool ships to everyone at once, the vendor contract covers the whole company, or the workflow does not divide cleanly. Three fallbacks are usable, and one thing you should stop doing.

Interrupted time series. Take a long pre-period, fit the trend, and test whether the post-launch series departs from the projected trend rather than from the last pre-launch point. This is materially stronger than a naive before-and-after because it survives a metric that was already improving, which is the single most common way AI gets credit for someone else's work. State the assumption out loud: nothing else changed at the same moment.

Non-equivalent comparison series. Find a metric that should not respond to the AI but should respond to everything else affecting the business, and track it alongside. If handle time drops in the AI-assisted queue and stays flat in the untouched queue over the same weeks, seasonality and staffing are less plausible explanations. This is weaker than a control group and much better than nothing.

Dose-response. If adoption varies across teams, and it always does, check whether the size of the effect tracks the intensity of use. A real effect usually scales with dose. A result that is identical for heavy users and non-users is telling you something, and it is not that the AI worked.

Stop converting hours saved directly into dollars. This is the most common inflation in AI business cases and the one the evidence undercuts most directly. Humlum and Vestergaard measured 3% time savings among Danish workers and no detectable change in earnings or hours, because the freed capacity went into new integration and oversight tasks. If your business case books saved hours as money, you owe the reader a sentence explaining what happened to that capacity: headcount not hired, overtime not paid, or volume absorbed without adding people. If none of those is true, the hours are real and the money is not, and the honest presentation says so.

When you cannot get a counterfactual at all, the right move is to label the claim accurately rather than dress it up. "Handle time fell 22% over the twelve weeks after launch, against a flat control queue" is a strong claim. "Handle time fell 22% after launch, with no control available" is a weaker one, honestly stated, and it will survive the first skeptical question. The inflated version will not.

Why AI ROI disappoints: the pilot trap

AI ROI disappoints for reasons that are almost never the model, and the published failure analyses converge on a short list.

Gartner's I&O survey is the most specific public evidence on why AI ROI fails to show up. Among leaders reporting at least one failure, 57% said their initiatives failed because they expected too much, too fast, assuming AI would immediately automate complex tasks or fix long-standing operational problems. Skills gaps and data problems each accounted for 38%. Notably, most AI success in that data comes from generative AI applied to IT service management and cloud operations, where the markets are mature: 53% of reported wins were in ITSM. Success clusters in the boring, well-understood places.

  • The pilot trap. A demo is cheap to produce now, which creates the illusion that the hard part is done. The pilot proves the AI can work once, on data you chose, in front of a friendly audience. The return only exists in production, on data you did not anticipate, at a volume and reliability the pilot never had to reach. S&P Global's finding that organizations scrap 46% of projects between proof of concept and broad adoption is the size of that gap, measured.
  • No baseline, so no provable return. The project never measured the before, so even when the AI helps, nobody can prove it, and an unprovable return gets treated as no return by the people holding the budget.
  • The denominator was wrong. The business case counted the model and ignored data, integration, evaluation, governance, change management, and inference. The real cost arrives later and the ratio collapses.
  • Inference quietly ate the margin. The per-prediction cost that nobody budgeted erodes a thin per-transaction return until the thing costs more to run than it saves. It shows up months after launch, when the enthusiasm has moved on.
  • Nobody redesigned the work. The AI was layered onto an unchanged process, so it produced a marginal improvement instead of a step change. This is the difference between McKinsey's 6% and everyone else, and it is a choice about how much you are willing to change.
  • The problem was never sharp enough to have a return. Gartner expects more than 40% of agentic AI projects to be cancelled by the end of 2027, citing escalating costs, unclear business value, and inadequate risk controls. Worth noting how thin the underlying data is: the supporting poll was 3,412 webinar attendees in January 2025, and the 40% is a prediction rather than a measurement. Gartner also estimates that only about 130 of the thousands of vendors marketing agentic AI are the real thing, which is a market-structure claim worth more than the headline percentage.

Notice what is absent from that list. The model was rarely the problem. The return disappoints because of the work around the model, the measurement nobody set up, and the redesign nobody did. All of those are fixable with discipline rather than with a bigger budget or a better foundation model. If you are weighing whether a specific use case will pay off before you commit, our AI development team scopes problems with exactly that lens.

What we still do not know about AI ROI

The honest inventory of AI ROI unknowns is short, important, and almost never published alongside the confident numbers.

Whether firm-level returns exist at scale. Every large positive figure in this article is self-reported. Every measured study is either narrow, like METR's 16 developers, or finds effects far smaller than the survey figures, like the 3% time savings in the Danish data. No study yet links verified AI deployment to audited financial statements across a large sample of firms. Until one does, the honest position is that firm-level AI return is plausible, widely believed, and not yet demonstrated with the rigor we would demand of any other capital allocation.

Whether time savings convert to money, and how long it takes. The Danish result rules out earnings effects above 1% two years in, while the authors expect a productivity J-curve that eventually turns up. Both of those can be true. Nobody knows the shape of the curve, and anyone who quotes you a payback period is quoting a plan, not a measurement.

Whether the developer productivity result generalizes. METR's finding is the strongest measured evidence that AI can be net negative for skilled workers on familiar ground, and it comes from 16 people using early-2025 tools. METR itself has since revised its experimental design. The direction of the effect on experienced engineers using current tools on current workflows is genuinely open. Our read on the practical state of those tools is in what AI coding tools do and do not do well.

Whether the quality cost is being counted. The Quarterly Journal of Economics study found small quality declines among the most experienced agents even as speed improved. Almost no corporate AI ROI calculation has a quality term in it. If AI trades quality for throughput in your workflow and you only measure throughput, your AI ROI number is measuring half the trade.

What the base rate is now. Almost everything above was fielded between mid-2024 and late 2025, on model generations that have since been replaced. The most-cited failure statistic in the field has a primary document that is no longer public. Treat every number here as a snapshot with a date attached, which is why each one carries its date.

None of this argues for waiting. It argues for measuring your own case rather than importing someone else's average, because the averages are describing populations you are probably not in. The organizations getting real AI ROI are not luckier: they scoped a sharp problem, counted the real cost, captured a baseline, rebuilt the work around the tool, and designed the rollout so the comparison would be valid. That is the whole difference, and none of it requires a better model.

Build an AI use case with a return you can actually defend
Talk to our AI engineers

Frequently asked questions

  • There is no defensible industry benchmark, because the published numbers measure different things. The most-quoted positive figure, 3.70 dollars returned per dollar invested, comes from an IDC InfoBrief sponsored by Microsoft in November 2024, based on self-reported answers from more than 4,000 business leaders. The most-quoted negative figure, a 95% pilot failure rate, comes from an MIT Project NANDA report whose PDF is no longer served at its official address. A useful working answer: a good AI ROI is one you can trace to a baseline you captured before the build and defend against a plausible alternative explanation. A measured 1.6x on an instrumented use case is worth more to a board than a claimed 10x nobody can reconstruct.

  • Three structural reasons, none of them technical. Attribution: AI usually arrives alongside a process change, a reorg, or a pricing move, and most teams never captured the baseline that would let them separate the effects. Cost visibility: the model API is the cheap, visible part, while data work, integration, evaluation, governance, change management, and recurring inference are where the money goes. Indirect value: much of the benefit lands as time saved rather than as a line on the income statement, and time saved does not convert to money automatically. A Gartner survey of 227 chief sales officers conducted in August and September 2025 found 31% naming difficulty proving the ROI of AI-driven tools as a top challenge for 2026.

  • Anywhere from 20% to 95%, depending entirely on what the counter counted. Gartner's April 2026 survey of 782 infrastructure and operations leaders found 20% of AI use cases fail outright and 28% fully succeed and meet ROI expectations, leaving a 52% middle band that is neither. S&P Global Market Intelligence found 42% of companies abandoned most AI initiatives before production in 2025. McKinsey's November 2025 survey found 61% of organizations attribute no enterprise-level EBIT impact to AI at all. The MIT NANDA figure of 95% counts pilots, not companies, and treats anything short of measurable P&L impact as failure. These are not contradictory findings. They are different questions.

  • Net value created divided by fully loaded cost, over a defined period, against a baseline you captured before the build. The numerator is the measured change in your chosen metric converted to money: hours saved times loaded cost, avoided loss, or attributable revenue. The denominator includes data work, integration, evaluation, governance, change management, and the recurring inference bill, not just the model. The part that makes the calculation credible is the counterfactual: an A/B test, a holdout group, or a staggered rollout that lets you say what would have happened without the AI. Without that, you have a before-and-after, which is a weaker claim and should be labeled as one.

  • Far more than the model. The visible cost is the API call or the licence. The real cost is sourcing and cleaning data, integrating into live systems, building an evaluation harness so you can prove the output is good, standing up governance and access controls, the change management that gets people to use the tool, and the recurring inference bill that charges you for every prediction after launch. In a serious implementation the model is often the smallest line item. Gartner's July 2024 prediction that at least 30% of generative AI projects would be abandoned after proof of concept by the end of 2025 named escalating costs and unclear business value among the causes.

  • From four places, in rough order of how reliably teams realize them: cost and time savings on high-volume repetitive work, revenue lift in marketing, sales, and product development, risk and error reduction in fraud and compliance and quality control, and new capability that was previously out of reach. McKinsey's November 2025 survey of 1,993 executives across 105 nations found cost benefits concentrated in software engineering, manufacturing, and IT, and revenue increases concentrated in marketing and sales, strategy and corporate finance, and product and service development. The same survey found that fundamentally redesigning workflows was among the strongest contributors to measurable business impact of all factors tested.

  • They agree on the mechanism and disagree on the sign, because they studied different people doing different work. Brynjolfsson, Li, and Raymond, publishing in the Quarterly Journal of Economics in 2025, measured 5,172 customer-support agents and found a 15% average increase in issues resolved per hour, with the largest gains among the least experienced workers and small quality declines among the most experienced. METR's randomized controlled trial, published in July 2025, found 16 experienced open-source developers were 19% slower on 246 tasks in repositories they knew well. Both results are credible. Skill level and task familiarity flip the sign.

  • The IDC InfoBrief sponsored by Microsoft reported organizations realizing value within 13 months and deployments taking under eight months, but those are self-reported figures from a vendor-sponsored study and should be read as an optimistic bound. Slower answers exist in the measured literature. Humlum and Vestergaard, studying roughly 25,000 Danish workers across 11 AI-exposed occupations linked to administrative records, found 3% time savings and no detectable effect on earnings or hours two years in. Set your own payback horizon before you start, and remember that recurring inference cost keeps moving the line after launch.