Frequent reports about how companies with AI-led processes are outperforming their peers is helping drive a huge interest in introducing Generative AI into core business workflows. As revealed in surveys like this one by Reuters, even smaller firms in sectors with high regulatory or data sensitivity like legal services and finance are increasingly adopting GenAI in one or more business functions.
In parallel, larger organizations are quickly moving beyond the traditional turn-by-turn conversation of a chatbot and adopting AI agents that autonomously plan and chain together multiple steps to achieve various goals, from booking travel to completing research that requires logging into accounts, running code, and compiling results into spreadsheets or slides.
According to the KPMG’s latest AI Quarterly Pulse Survey conducted between May and June of 2025 with “130 top-tier U.S.-based executives and business leaders, all from organizations boasting annual revenues of $1 billion or more,”
Organizations are rapidly accelerating from experimentation to piloting AI agents.
Still, according to the Preliminary Findings from AI Implementation Research from Project NANDA, despite $30–40 billion in enterprise investment into GenAI, 95% of organizations are getting zero return.

Examples abound of failed trials, like the one by the UK government using Microsoft Copilot that resulted in “no discernible gain in productivity.”
What’s the secret to extracting real value out of GenAI?
First, to repeat the requirements you’ll find in most articles written by consulting firms or AI leaders describing how to move the needle with GenAI:
Successful change management, with “mobilization of the C-suite leadership to effectively drive AI adoption and scaling.”
Blended teams that integrate specialized external talent with full-time employees.
An “assemble approach” that customizes solutions for specific business needs based on open-source building blocks that can be easily updated or swapped out (a key step when “the shelf life for state of the art AI is shorter than a jar of organic marinara sauce”).
To understand why all these elements may not suffice, let’s at a fictionalized case study in a domain in which I’ve spent two decades working: software development.
Company B hears about the excellent results a competitor, Company A, achieved using AI-assisted software development to increase the productivity of their developers. It hires a consulting firm to help with GenAI capability building, redesigns its software development process from the ground up, and establishes robust change management mechanisms to prevent resistance or training gaps. But despite doing everything “right” according to expert advice, at the end it can’t reproduce any of the productivity gains claimed by Company A.
What went wrong? Assuming that Company A wasn’t exaggerating its results and execution in Company B wasn’t flawed, the missing element is likely to be untested assumptions that led one company to copy what was done in another without understanding why it worked.
Clearly, this is not a problem limited to investments in GenAI (thus the “beyond” in the title reflecting the fact that the lessons here are applicable broadly to significant investments). At the core of the issue is the lack of a reliable foundation for making informed investment decisions.
Misunderstanding the “why” may lead to weak (or negative) results
In our example, the development teams in both Company A and Company B create software primarily in Java and React. However, in Company A, the developers were being slowed down by labor-intensive and easier-to-inspect tasks like documentation and test case generation (things the AI assistant solution excels at). In contrast, at Company B the bulk of the work is about optimization, defect fixing, and hard-to-inspect tasks like verifying system architectures. No one with knowledge of vehicles would assume that a car that performs well on smooth roads would necessarily work off-road, so it shouldn’t be a surprise that the solution that yielded productivity gains for Company A failed in Company B.
Humans are terrible predictors
There are numerous studies in behavioral science, psychology, and decision theory that explain why humans are generally poor at predicting the future, especially in complex, uncertain environments. With that in mind, it makes little sense to pay attention to what people '“expect” will become true in the next months or years, or what outcomes they “believe” will be obtained from an investment in technology or workflow redesign.
Take for example the results of a randomized controlled trial (RTC) used “to understand how AI tools at the February–June 2025 frontier affect the productivity of experienced open-source developers.”
In this experiment, the developers displayed an overoptimistic opinion of how AI affects their productivity, both before and after completing tasks. Before starting work, developers forecast that AI would reduce issue completion time by 24%. After the work, their estimate was a bit lower, 20% on average. In reality, the study measured a negative effect of AI assistance on their productivity.
The Bottom Line: Evidence-based decisions are the only reliable foundation for positive ROI in any significant investment
The saddest part of the current state of affairs is that many AI-assisted workflows that could achieve persistent value will be abandoned after millions are wasted in failed pilots or flawed implementations.
What’s missing for the organizations on the wrong side of the GenAI divide (high adoption, low transformation) is an evidence-based approach to investment decisions.
In its modern form, the “evidence-based approach” was popularized by the influential work of Professor Archie Cochrane, who in 1972 argued for the need for systematic reviews of clinical evidence, changing medical practices.
Applied to investments in digital business transformation, an evidence-based approach requires going beyond the veneer of credibility of vendors and their successful case studies. It calls for framing your assumptions as testable questions. It demands independent research that isn’t tainted by conflicts of interest or weak claims that only look like solid evidence. It avoids common pitfalls such as survivorship bias, where failed projects are excluded from the analysis, or inappropriate benchmarks that don’t reflect the context of how a solution will be used.
Lack of access to formal studies or reported data is not a valid excuse for failing to adopt an evidence-based approach to business innovation. Smart organizations protect their large investments by first identifying all relevant assumptions and creating testable hypotheses (“We believe that implementing [new tool/process] will reduce [X] by [Y]% in [Z] time.”). And from there, they use experiments, data, and feedback to develop their own evidence in a systematic, rigorous way.



