Back to Blogs

The AI Execution Gap: Why Most AI Projects Fail After Adoption

August 14, 2026 / 24 min read / by Irfan Ahmad

The AI Execution Gap: Why Most AI Projects Fail After Adoption

Share this blog

AI is spreading fast across companies, but real value is not. The break usually comes after the pilot, when promising systems run into weak data, unclear ownership, and un-reworked workflows. It is turning that capability into something reliable, governed, and economically useful at scale.

Where AI Starts Colliding with Reality

Where AI Starts Colliding with Reality

In February 2024, Klarna released the kind of AI story every executive wanted to hear. Its AI assistant, built with OpenAI, had handled 2.3 million conversations in its first month, accounted for two-thirds of the company’s customer-service chats, performed work equivalent to roughly 700 full-time agents, reduced repeat inquiries by 25 percent, and cut average resolution time from 11 minutes to less than 2.

Klarna said the impact could contribute roughly $40 million in profit improvement during 2024. It was a clean, persuasive case for enterprise AI. A large consumer-facing company had deployed AI in a real workflow, attached clear operating metrics to it, and presented the result as proof that the leap from experimentation to execution was already underway.

That is exactly why the next part matters. A few months later, McDonald’s ended its AI drive-thru ordering pilot with IBM after testing the system in more than 100 US restaurants. The use case had looked ideal on paper. High volume, repetitive demand, measurable throughput, and clear labor economics. Yet the live environment turned out to be less obedient than the model logic behind it.

Voice ordering in a drive-thru is not just speech recognition. It is accents, speed, interruptions, changing orders, ambient noise, edge cases, and real-time service pressure. McDonald’s did not abandon AI forever, but the decision to end the pilot exposed something the market still struggles to say plainly.

A system can look highly credible in rollout decks and still fail to hold inside the operating chaos of real business. Klarna and McDonald’s are not simply a success story or a failure story. Both show that AI outcomes depend less on access to the technology than on how well the surrounding workflows, controls and human roles are redesigned around it.

That tension sits at the heart of enterprise AI in 2026. Adoption is spreading quickly, but value is not spreading in the same way. McKinsey’s 2025 global survey found that organizations are using AI widely, yet most are still far from enterprise-wide impact, and the firms seeing stronger bottom-line results are distinguished less by access to AI than by workflow redesign and senior-level oversight.

NBER researchers, looking at nationally representative US survey data, found that by late 2024 generative AI had already reached 39 to 45 percent of adults aged 18 to 64 in some form, with roughly a quarter of employed respondents using it for work in the previous week.

Gartner added a sharper warning. It first predicted in mid-2024 that at least 30 percent of generative AI projects would be abandoned after proof of concept by the end of 2025, then said in early 2026 that the share had reached at least 50 percent.

The issue is no longer whether AI can produce something useful. In many cases, it clearly can. The issue is whether organizations can turn that usefulness into something repeatable, governable, and economically durable once the technology leaves the demo and enters the company’s real operating core.

AI can produce something useful

This is the execution gap. It is the space between a capability that works in isolation and an institution that knows how to make that capability hold under pressure. It appears when a model performs well, but the workflow around it does not work. It widens when data is fragmented, ownership is split, review burdens rise, and economic value becomes harder to prove than technical promise.

It is now becoming one of the most important divides in the AI economy. The pattern in enterprise AI in 2026 is now hard to ignore. Many AI projects fail because usefulness in isolation is not the same thing as execution in a real business. The next phase of AI will not be decided only by who builds the best models. It will be decided by those who can make these models thrive in real-world conditions.

Where the Gap Becomes Visible: The Split Between AI Use and AI Value

The execution gap becomes visible at the point where AI adoption starts to look broad, while enterprise value still looks narrow. On the surface, the market appears to be moving quickly. McKinsey’s 2025 global survey found that 78 percent of respondents said their organizations use AI in at least one business function, up from 72 percent in early 2024 and 55 percent a year earlier, which is a strong signal that AI has already moved into mainstream business use.

NBER’s nationally representative research points in the same direction. By late 2024, nearly 40 percent of Americans aged 18 to 64 were already using generative AI, while 23 percent of employed respondents had used it for work in the previous week and 9 percent used it every workday. These numbers show a technology that has already entered ordinary organizational life.

Computer, Internet and AI Adoption

Yet the value side of the picture is much harder. The same McKinsey research shows that AI has still not significantly affected enterprise-wide EBIT for most organizations, and even among respondents who reported some EBIT impact, the majority said less than 5 percent of EBIT was attributable to AI use. This is a very different story from the surface narrative of rapid transformation.

BCG’s October 2024 findings make the gap sharper still. In one release, BCG said 74 percent of companies still struggle to achieve and scale value from AI, while in a related analysis it said only 22 percent of companies had advanced beyond proof of concept and only 4 percent were creating substantial value. These numbers are worth sitting with because they show how easy it is for adoption to outrun conversion and highlight how the market is still short on scaled outcomes.

Deloitte’s enterprise research helps explain why the contradiction persists. In its Q4 report, more than two-thirds of respondents said that 30 percent or fewer of their generative AI experiments would be fully scaled in the next three to six months, a finding Deloitte used to argue that however quickly the technology advances, organizational change inside large firms still move much more slowly.

A company can be enthusiastic, funded, active, and publicly vocal about AI, while privately expecting that most of its experiments will not become scaled operating systems anytime soon.

This is where the real divide now sits. McKinsey’s report points in that direction very clearly. Among the attributes it tested, workflow redesign had the strongest effect on an organization’s ability to see bottom-line impact from generative AI.

This matters because it suggests that the hard part begins after the model works. Once the model is credible enough to use, the question shifts to whether the business around it changes enough for the value to hold. That is the first place where the execution gap becomes visible.

Why Working AI Still Fails: Difference Between Model Success and System Success

The easiest way to misunderstand enterprise AI is to assume that once a model performs well, the hard part is over. In practice, the harder part often starts there. A pilot proves that a system can do something useful under managed conditions.

Production asks a different question. Can the same system keep doing useful work when the inputs get messy, when customers behave unpredictably, when policies collide, when liability appears, and when the output must survive in contact with a real workflow rather than a demo environment? This is where many projects begin to weaken.

S&P Global Market Intelligence reported in 2025 that the average organization scraps 46 percent of its AI proof-of-concept projects before production, and that the share of companies abandoning most of their AI initiatives before production had risen from 17 percent to 42 percent year over year.

One reason is that AI often fails in live environments that look simple from a strategy deck. New York City’s MyCity chatbot was launched to help business owners get official guidance from city information pages. It was pitched as a practical public-service use of AI.

Yet reporting by The Markup found the bot telling employers they could take workers’ tips, telling landlords they could reject some tenants using housing vouchers, and telling businesses they could refuse cash, answers that were not merely inaccurate but, in some cases, unlawful as well. The city kept the tool live while adding stronger warnings that the system could produce “inaccurate or incomplete” responses.

Reuters later reported that New York City’s then Mayor Eric Adams defended the pilot even after the errors became public: Fast Company / Reuters – NYC mayor defends chatbot pilot. The problem was that a system placed inside a legal and operational workflow had to be right in ways a general-purpose model is often not.

Chevrolet dealership chatbot went viral

A second pattern appears when the model is responsive, fluent, and still impossible to contain. In late 2023, a Chevrolet dealership chatbot went viral after users prompted it to do things far outside its intended role. It recommended a Ford F-150 over Chevy vehicles, answered unrelated coding questions, and in one widely shared exchange agreed that a 2024 Chevrolet Tahoe could be sold for $1.

The point of the story was not that anyone actually bought the car for a dollar. It was that a customer-facing AI layer had been put in front of a commercial workflow without meaningful boundaries. Once users realized the system could be steered, the dealership lost control of the interaction almost immediately. This is a small example, but a revealing one. The model did not need to be terrible to fail as it only needed to be insufficiently constrained in a real commercial setting.

A third pattern emerges when performance risk turns into legal responsibility. In the Air Canada bereavement-fare case, a customer relied on incorrect information given by the airline’s chatbot about retroactive fare refunds after a family death.

When the dispute reached a British Columbia tribunal, Air Canada argued that the chatbot was a “separate legal entity” responsible for its own words. The tribunal rejected that claim and held the airline liable. The lesson is larger than travel support. Once AI enters a customer-facing workflow, responsibility does not dissolve into the software. The organization still owns the output, the consequences, and the trust damage when the system is wrong.

Why Old Workflows Keep Breaking New AI: The Process Problem Inside Firms

Most companies are still trying to fit AI into workflows built for earlier systems. This is where the execution gap starts to widen. Traditional software usually sat inside a defined process where inputs were structured, outputs were predictable enough, and responsibility was easier to assign. AI changes the shape of the work itself. It introduces probabilistic outputs, new review layers, new exception paths, and a heavier dependence on data quality, context, and judgment.

McKinsey’s 2025 global survey puts the point bluntly. Of the 25 attributes it tested, workflow redesign had the biggest effect on an organization’s ability to see EBIT impact from generative AI. Yet only 21 percent of respondents reporting gen AI use said their organizations had fundamentally redesigned at least some workflows as they deployed it.

AI Risk Management Framework

This split matters because it explains why AI can be everywhere in a company and still remain shallow. McKinsey’s separate workplace report from early 2025 found that almost all companies invest in AI, but only 1 percent believe they are at maturity.

The same report argues that the biggest barrier is not employee willingness to use AI, but leadership’s failure to steer the organization fast enough. This gap between tool access and operating redesign is where many firms now sit. They have copilots, assistants, and pilots in circulation, but the core workflow still assumes human handoffs, fragmented systems, and old approval logic.

The companies getting further usually look less like fast adopters and more like careful process engineers. Morgan Stanley is one of the cleaner examples. Its AI assistant for wealth management was not treated as a loose chatbot experiment.

The bank trained it on more than 100,000 research reports and documents, held back broad deployment until evaluation frameworks showed the answers met advisers’ quality standards, and then built follow-on tools such as AI @ Morgan Stanley Debrief, which turns client meetings into summaries, action items, draft emails, and notes saved into Salesforce.

OpenAI says over 98 percent of Morgan Stanley adviser teams now use the assistant daily, while document access jumped from 20 percent to 80 percent. These numbers matter, but the more important point is structural. Morgan Stanley did not just add AI to the existing workflow. It reshaped the flow of research retrieval, meeting follow-up, and adviser support around it.

The broader research points in the same direction. BCG reported in June 2025 that employees at organizations undergoing comprehensive AI-driven redesign were experiencing a visibly different workplace from those at less advanced firms, which is another way of saying value tends to show up where the work itself changes, not where AI stays confined to scattered tools.

AI Workflow Redesign

Deloitte’s 2026 enterprise report uses a similar frame. It separates firms using AI to deeply transform core processes, business models, or products from those still using it more superficially, with limited change to the underlying process.

This is a useful distinction because it reflects the real divide. The market is no longer separating users from non-users. It is separating firms that have started to rebuild how work moves from firms that are still layering AI onto old process maps.

Why Enterprise Data Still Fails AI: The Data and Context Trap

A great deal of enterprise AI weakens at the point where the model meets the company’s actual information environment. Firms often assume they have the raw material for AI because they have data, documents, dashboards, wikis, customer histories, and years of accumulated reporting.

In practice, much of that material is fragmented, duplicated, stale, trapped behind permissions, or stripped of the business context that made it useful to humans in the first place.

IBM’s 2025 survey of AI adoption challenges captures the problem cleanly. Forty-five percent of respondents cited concerns about data accuracy or bias, 42 percent said they lacked enough proprietary data to customize models, 42 percent pointed to inadequate generative AI expertise, and 40 percent cited privacy or confidentiality risks.

Gartner has made the same point more bluntly. By the end of last year, it said, at least 50 percent of generative AI projects had been abandoned after proof of concept, with poor data quality among the main reasons.

The problem is not only bad data. Missing context also plays a significant role. A model can read a document, but that does not mean it understands which version is authoritative, which exception still matters, which internal term means something different in one business unit than another, or which answer is acceptable in a regulated setting. This is why unstructured information keeps turning into an execution bottleneck.

Box’s 2025 note on structured vs. unstructured data in the age of AI is useful here because it makes a point many firms still underestimate: most enterprise knowledge does not live in neat rows and columns at all. It lives in contracts, PDFs, call notes, decks, policy docs, emails, chats, and uploaded files, which makes retrieval, extraction, and permission-ing far harder than the sales pitch around “enterprise knowledge” usually implies.

MIT Sloan makes a related observation from a more operational angle. In a February 2026 piece on AI implementation in finance, it notes that 60 percent to 80 percent of time in a data analytics project may be spent acquiring and cleaning data before the more advanced AI work can even begin.

This is why some of the more convincing enterprise examples are really stories about context engineering. BBVA’s legal-services chatbot in Mexico did not work because a general model suddenly became a banking lawyer. It worked because the bank constrained the system to standardized, pre-validated legal FAQs and documentation guidance reviewed by its own Legal Services team.

According to OpenAI’s 2025 enterprise report, the tool now automates more than 9,000 queries annually, has enabled BBVA to redeploy the equivalent of 3 FTEs, and contributes 26 percent of the Legal Services division’s annual savings KPI. The underlying lesson is not that AI solved legal ambiguity on its own. It is that approved internal knowledge, tightly bounded to a real workflow, became usable at speed.

Oscar Health shows the same pattern in a more consumer-facing environment. OpenAI’s report says Oscar’s platform answers 58 percent of benefits questions instantly and handles 39 percent of benefits messages without human escalation. Oscar’s own 2025 write-up on Oswell makes an important point.

The agent is useful because it is connected to claims, medical records, prior care interactions, and benefit logic inside Oscar’s own systems, while still routing people to human teams when needed. Context is doing the real work there as the system is grounded in member data, plan rules, and service pathways that a general model on its own would not reliably infer.

Moderna offers a third version of the same lesson, but inside a scientific organization rather than a consumer workflow. In its April 2024 note on its collaboration with OpenAI, Moderna said it had already launched mChat, its own instance of ChatGPT built on OpenAI’s API, before extending deployment through ChatGPT Enterprise across clinical development, legal, and corporate functions.

The company’s own framing matters. It does not describe AI as a standalone magic layer. It describes a broader data platform, personalized support, and real-time training that help employees use AI against Moderna’s own internal context.

OpenAI and Moderna

OpenAI’s enterprise report then gives a more specific downstream example, saying Moderna used AI to compress a core analytical step in target product profile planning from weeks to hours in some cases. The pattern is consistent. The gain does not come from generic intelligence floating above the business. It comes from a model meeting an internal knowledge environment that has been made usable enough to support a specific decision process.

Who Actually Owns AI Within the Firm: The Ownership Vacuum

Many AI projects slow down or drift once they leave the pilot phase because nobody fully owns them in the phase that matters most. In a pilot, ownership looks simple enough. Once the system enters real work, that arrangement starts to break. The output affects operations, compliance, customer experience, legal exposure, and internal decision-making at the same time.

McKinsey’s 2025 global survey gets close to the heart of the problem. It found that CEO oversight of AI governance is one of the elements most correlated with higher self-reported bottom-line impact from generative AI, yet only 28 percent of respondents said their CEO was responsible for overseeing AI governance, while 17 percent said the board held that responsibility.

In many companies, McKinsey noted, governance is jointly owned, with respondents reporting that two leaders are in charge on average. This sounds collaborative but in practice, it often means accountability is spread across enough people for no single one to truly carry it.

The deeper issue is that AI changes where decisions are made and who bears the consequences when those decisions go wrong. The World Economic Forum’s 2026 paper on organizational transformation argues that the challenge is no longer whether AI works, but how organizations must re-architect workflows, operating models, and decision rights so AI can be embedded into execution.

It makes the point even more clearly later in the report, where it says organizations advancing beyond experimentation are converging on a few shared principles, including clear business ownership of AI, workflow redesign rather than pilot expansion, and human accountability at scale.

In the Forum’s phrasing, the shift required is from human-in-the-loop to human-in-the-lead, with explicit decision ownership, autonomy thresholds, and escalation paths defined before, during, and after deployment at scale.

That distinction matters because firms often confuse oversight with ownership. Oversight means someone reviews the system whereas ownership means someone is responsible for the workflow, the boundaries, the escalation logic, the metrics, the exceptions, and the consequences.

IBM’s February 2026 implementation guide puts it in more operational terms. It argues that governance becomes real only when an organization defines people, roles, and authority across the AI lifecycle, with named responsibilities for model owners, business-unit leads, engineering teams, and risk officers.

Its language is not abstract. It talks about clear reporting lines, decision authority, risk registers, rollback processes, and real-time monitoring and accountability. This is a useful reminder because many firms still speak about responsible AI as a values statement when the harder issue is operating design.

AI Governance Framework

The more advanced company examples are starting to reflect that shift. The World Economic Forum’s 2026 paper points to Repsol, which redesigned parts of its operations around a human-in-the-loop, agentic AI model where custom agents execute discrete tasks, gather inputs, run checks, draft outputs, and trigger actions within guardrails, while humans retain control through review, approval, and exception handling.

The report says 22 agents are already live across 38 use cases, with plans to scale to more than 90 agents and support over 3,000 IT employees. What matters in that example is the operating principle. Repsol defined the boundary between machine action and human control inside the workflow itself.

Keeping human in loop

The same report gives another clue about where the organization is headed. It says advanced adopters are redesigning management roles around orchestration, judgement, system stewardship, and guardrail refinement, while workforce models begin to codify access rights, autonomy boundaries, escalation paths, and lifecycle governance because accountability becomes contested when agents act across workflows.

Ownership vacuum has become one of the decisive parts of the execution gap. The ownership vacuum is a structural lag between older management models and newer systems that act with partial autonomy. Companies can live with that lag during experimentation, but they struggle once AI starts touching execution.

Once a model starts influencing actions, recommendations, approvals, or customer outcomes, the question then is who owns the judgement around it, who can override it, who carries the risk when it fails, and who has the authority to redesign the workflow when the old structure stops working.

What Firms Getting Real Value Are Doing Differently

The companies getting meaningful value from AI tend to share a few traits. They start with specific business problems rather than broad AI mandates, redesign workflows instead of simply adding tools, keep humans involved where judgement still matters, measure outcomes rather than usage, and give senior leaders clear ownership of execution.

The next divide in enterprise AI is unlikely to sit between companies that use AI and companies that do not. A more important split is opening between firms that have folded AI into the structure of the business and firms that are still treating it as a layer of tools. Deloitte’s 2026 enterprise report captures that divide neatly.

Thirty-four percent of surveyed organizations say they are starting to use AI to deeply transform their business, whether by creating new products and services or reinventing core processes and business models. Another 30 percent say they are redesigning key processes around AI.

The remaining 37 percent are still using AI at a surface level, with little or no change to the underlying process. This shows where the real competitive line is moving. AI use is spreading widely, but deep organizational conversion is not.

BCG’s work makes the same point from another angle. Its June 2025 report says business value from AI requires deep workflow redesign, while AI usage by itself has already become mainstream. It also reports that three quarters of respondents believe AI agents will be vital for future success, yet only 13 percent say agents are currently integrated broadly into workflows. That gap is revealing. Many companies can imagine the next phase of AI.

Far fewer have built the operating conditions for it. BCG also found that companies reshaping workflows and functions with AI are producing a visibly different employee environment from those that are only rolling out off-the-shelf tools. This is a sign that value starts compounding when AI changes how work moves, and not merely how fast a few tasks get done.

The firms that get further usually share a few traits. They narrow their use cases early and redesign the workflow rather than simply inserting a model into an old one. They connect AI to approved internal knowledge and define ownership, review thresholds, and escalation logic before scale.

McKinsey’s 2025 survey has already shown that workflow redesign is the attribute most strongly linked to self-reported EBIT impact, while its workplace study says nearly every company is investing in AI, but only 1 percent believe they are at maturity. These numbers help explain why the winners will be firms with cleaner operating systems.

There is also a more practical reason the winners will begin to separate more clearly from the rest. Once AI capability becomes easier to access, the advantage shifts away from mere access and toward institutional translation. A model that any rival can license is not, by itself, a strategy.

A company that can convert this model into a faster decision cycle, a better service architecture, a lower-cost process, a safer review framework, and a better trained workforce begins to build something harder to copy.

BCG’s March 2026 analysis on AI-first cost advantage puts it more directly. It says AI leaders are seeing three times greater cost reduction, 1.6 times higher EBIT margins, and 2.7 times stronger total shareholder return than laggards. This is a useful metric because it shifts the conversation away from novelty and back toward operating performance.

So, the firms that win from AI are unlikely to be the loudest adopters. They will more often be the ones that look operationally different under the surface. The market has spent the last two years asking who is using AI. The more important question now is who is reorganizing its workflows around it.

Conclusion: The Real Test of Enterprise AI

The AI economy is defined through capability. Better models, cheaper inference, broader access, faster deployment, more agents, more automation. These shifts are real, and they explain why AI has moved so quickly from experimentation into everyday business use. They do not, however, explain where durable enterprise value will settle. The harder question now is execution.

McKinsey’s research has shown that AI use is already broad, but the firms reporting stronger bottom-line impact are distinguished by what they changed around it, especially workflow redesign and senior-level governance. The critical issue now is no longer who can buy intelligence, but who can reorganize the business so that intelligence can hold.

That is why the execution gap should not be treated as a temporary phase in adoption, or as a problem that will simply fade as models improve. It is becoming one of the main filters through which AI value is being distributed. Deloitte’s 2026 enterprise report shows that only a minority of firms are using AI to deeply transform core processes, business models, or products, while a large share still remains at a more surface level of use.

BCG reaches a similar conclusion from a different direction. Its 2025 and 2026 work suggests that firms pulling ahead are distinguished by the harder work beneath the surface: redesigning workflows, integrating agents into the operating model, tightening decision structures, and converting model access into cost, margin, and productivity effects that competitors cannot easily reproduce.

The central question has therefore changed. Earlier, the market wanted to know whether generative AI could be useful at all. That question has been answered strongly enough to move capital, strategy, and organizational attention. The question now is where that usefulness holds under pressure. The next divide in AI will not be as simple as frontier labs versus everyone else, or adopters versus holdouts. The companies that win will have institutions better designed to absorb, govern, and compound what these AI tools make possible.

Therefore, the AI execution gap is not an afterthought issue anymore. It is becoming one of the main mechanisms through which AI value is being calculated. The market has spent enough time asking who has access to intelligence. The more consequential question now is who can build an organization that knows how to use it well enough, consistently enough, and structurally enough for the value to last.