Generative AI in the enterprise: where the value is real, and where it's still a gadget
What randomized studies (MIT, NBER, METR) reveal about where generative AI creates real enterprise value, and where it remains a gadget.
Upleo
The figures circulating about generative AI in the enterprise seem to contradict each other: 88% of organizations already use it in at least one function, but 95% of pilot projects produce no measurable financial return. These two figures aren't incompatible — they describe two different things. Adoption is massive; transformation, on the other hand, remains rare. The real value of generative AI isn't uniform across all tasks: it's now demonstrated, through controlled studies, on structured, repetitive work with a clear success criterion — and far less established, even negative in some measured cases, on expert and creative work.
The 2026 paradox: near-universal adoption, vanishingly rare value
Three figures, published a few months apart by three independent organizations, sketch the same landscape from different angles.
The report The GenAI Divide: State of AI in Business 2025 from MIT Media Lab (NANDA project), built from 150 executive interviews, a survey of more than 350 employees and an analysis of 300 real deployments, established that despite $30 to $40 billion in enterprise investment, 95% of generative AI pilots produce no measurable impact on the bottom line. Only 5% of integrated pilots generate substantial value.
McKinsey, in its State of Organizations 2026 report (a survey of more than 10,000 respondents conducted between June and September 2025), reaches a similar conclusion through a different method: 88% of organizations are experimenting with AI, but 81% report no significant impact on their results. Only 1% of organizations describe their deployment as fully mature. A finer segmentation identifies about 6% of "high performer" organizations, which attribute more than 5% of their EBIT to AI — a figure that independently confirms the order of magnitude found by MIT.
BCG rounds out this picture with a structural data point: in its analysis of hundreds of enterprise deployments, about 10% of the value generated by AI comes from the algorithms themselves, 20% from the data and technology needed to implement them, and 70% from people and processes. This ratio, which BCG calls the 10-20-70 rule, quietly explains why model performance is almost never the limiting factor.
A question of method: what evidence actually counts?
Before going further, a methodological distinction is worth making, because it directly affects the reliability of what can be claimed. Most of the figures cited about AI in the enterprise come from self-reported surveys — executives or employees are asked to estimate a productivity gain or value creation. These surveys have the advantage of large samples, but they measure a perception, not a causal effect.
A much smaller handful of studies rely on randomized controlled trials (RCTs) — the reference method in medical research and experimental economics, where identical tasks are randomly assigned to a group with AI access and a group without, isolating the tool's actual effect. Two of these studies, published or peer-reviewed, offer the strongest data available to date on the question — and they point in opposite directions depending on the type of work studied.
Where the value is demonstrated: structured, repetitive work with a clear success criterion
The reference study on this front is by Erik Brynjolfsson (Stanford), Danielle Li (MIT Sloan) and Lindsey Raymond, published in the Quarterly Journal of Economics in 2025 after an initial NBER working paper in 2023. The authors leveraged the staggered rollout of a conversational AI assistant across 5,172 customer support agents at a Fortune 500 company. The gradual, non-simultaneous rollout created the conditions for a natural quasi-experiment.
The result: access to the tool increased average productivity, measured in cases resolved per hour, by 15%. But the average masks a marked heterogeneity — the least experienced and lowest-skilled agents improved by 34%, in both speed and quality of handling, while the most experienced agents gained almost nothing, and even saw a slight quality decline in their exchanges. The researchers interpret this result as a skill-leveling effect: AI spreads the practices of top-performing agents to less experienced ones, rather than uniformly boosting everyone.
This finding echoes MIT NANDA's take on the sectoral distribution of value: in its sample of 300 real deployments, back-office automation — case processing, data reconciliation, repetitive administrative tasks — produces the highest and most reliable returns, by reducing outsourcing costs and processing times. Yet, according to the same study, this is precisely the function that receives the least AI investment budget, with most spending concentrated on sales and marketing — the functions where, conversely, the measured return is lowest.
The common thread across these documented success cases is identifiable: a task broken into comparable units, an objective success criterion (a case resolved, data correctly processed), and a fast feedback loop that makes it possible to verify whether the AI's output is correct or not.
Where the promise falls apart: expert work and poorly framed automation
The most rigorous counter-example comes from METR (Model Evaluation and Threat Research), an independent research organization that ran a randomized controlled trial in 2025 on the impact of AI tools on experienced developers — this time on expert work, not structured customer-support tasks.
Sixteen experienced developers, with an average of five years contributing to mature open-source repositories they already knew well, completed 246 real tasks (bug fixes, features, refactors), each task randomly assigned to a condition with or without access to AI tools. Result: developers with AI access took 19% longer to complete their tasks than those without access. Even more striking, those same developers, once the work was done, estimated on average that they had been 20% faster thanks to AI — a nearly 40-point gap between perception and actually measured time.
METR notes that this result applies to a specific context (experienced developers, repositories they already knew well, early-2025 tools) and that a follow-up published in February 2026 shows signs of improvement with more recent tools — but the perception gap itself stands as an independent signal from the generation of tools tested: self-reported AI productivity gains aren't reliable, even among experienced, good-faith users.
The same pattern of disappointed expectations shows up in agentic AI. Gartner predicts that more than 40% of agentic AI projects will be abandoned by the end of 2027, not for reasons of model performance, but because of unforeseen costs, poorly defined business value and insufficient governance. The firm estimates that barely 130 of the thousands of vendors claiming to be "agentic" offer genuinely autonomous capability — the rest falling under what it calls "agent washing," a marketing repositioning of existing automation tools or chatbots.
The common thread across these documented failures is also identifiable: unstructured work already mastered by an expert, with no unambiguous success criterion, or automation bolted onto an existing process without rethinking it.
The factor that explains the gap: it's not the model, it's the organization
MIT NANDA calls this phenomenon the "learning gap": an organization's inability to integrate an AI tool into its workflows, structures and culture — the friction isn't technological, it's human and organizational. The authors explicitly note that the divide between the 5% who succeed and the 95% who fail isn't determined by model quality or regulation, but by the deployment method.
McKinsey puts a number on this intuition. In a 2026 survey, organizational readiness — redesigned workflows, flexible technical platforms, clear governance — explains 48% of the captured-value gap between organizations, versus 25% for individual employee readiness. In other words, an individual employee's personal enthusiasm for AI matters half as much as the organization's ability to restructure its processes around the tool. The same report finds that companies that have redesigned their workflows, even at an early stage of adoption, are 5.3 times more likely to report measurable enterprise value creation than those that left their processes unchanged.
This reading aligns with BCG's 10-20-70 rule mentioned above: the technology itself represents only a minority fraction of total value. Organizations BCG classifies as AI leaders show cost reductions three times higher, EBIT margins 1.6 times higher and return on invested capital 2.7 times higher than their peers — a performance gap that isn't explained by differential access to models, available in almost identical form to everyone, but by how the work was rebuilt around them.
A checklist before launching a pilot
The studies converge on a small number of criteria that distinguish a high-potential use case from an expensive gadget:
- Is the task structured and repetitive, with comparable units of work from one case to the next — or is it expert judgment that's hard to standardize?
- Is there an objective, quickly verifiable success criterion — a right or wrong result, a case resolved or not — or does quality depend on a subjective judgment that's hard to settle?
- Will the process actually be redesigned, or is AI simply being added on top of an existing workflow with nothing else changing?
- Does the organization have proprietary data or context that differentiates its use of the tool from a competitor using the same generic model?
- Does a feedback loop exist so the system, or the teams using it, can learn from mistakes over time?
A project that checks most of these boxes resembles the documented success cases from Brynjolfsson, Li and Raymond, or from MIT NANDA on back-office work. A project that checks none of them resembles the 95% of pilots with no measurable return — or, in the case of poorly framed agentic AI, the 40% of projects Gartner expects to be canceled by 2027.
What this concretely changes for a company getting started
The most common trap isn't underinvesting in AI, but adopting it as a generic reflex — installing it without redefining the process, the success criterion, or the data that feeds it. That is, at bottom, the same mistake described in our article on brand positioning: changing the surface (a slogan, or here a tool) without changing what's behind it produces neither differentiation nor measurable value. Generative AI doesn't escape that rule — it simply illustrates it at a larger scale, and with more rigorous data than most enterprise transformation topics.
Sources
- Brynjolfsson, E., Li, D., & Raymond, L. (2025). Generative AI at Work. The Quarterly Journal of Economics, 140(2), 889-942. academic.oup.com
- Becker, J., Rush, N., Barnes, E., & Rein, D. (2025). Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. METR. metr.org
- MIT NANDA / MIT Media Lab (2025). The GenAI Divide: State of AI in Business 2025.
- McKinsey & Company (2026). The State of Organizations 2026. mckinsey.com
- Boston Consulting Group (2026). How Leaders Build an AI-First Cost Advantage; AI Transformation Is a Workforce Transformation. bcg.com
- Gartner, Inc. (2025). Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027. gartner.com
- Stanford Institute for Human-Centered AI (2026). The 2026 AI Index Report — Economy. hai.stanford.edu
Wondering whether an AI use case has real potential in your organization, or risks joining the 95% with no measurable return? Let's talk about it, starting from your actual context, not a generic demo.
Frequently asked questions
Related articles

From generic offer to defensible positioning: the method for finding what actually makes you different
How to turn an offer perceived as a commodity into a defensible positioning, with a concrete 3-step method and a real case.

Lead scoring and qualification: what automation actually changes in commercial prioritization
How to build a reliable B2B account score: a 4-pillar method, a step-by-step worked example, and calibration pitfalls to avoid.

Building an AI-assisted content pipeline without losing your brand voice
What a study published in PNAS Nexus reveals about homogenization by generative AI, and how to build a content pipeline that preserves a distinct brand voice.

Data governance: the invisible prerequisite before automating anything
What the RAND Corporation study reveals about the causes of AI project failure, and how to assess your data governance before automating.
A question about this article?
Let's talk about your context and how we can help.