Skip to main content
The JournalAug 17, 2026
Operations

Why 95% of AI Pilots Fail (And What to Automate First)

MIT found that 95 percent of corporate AI pilots produce no measurable return. Not because the models are weak, but because companies automate the wrong things in the wrong order. Here is the boring, unglamorous framework that actually ships.

Why 95% of AI Pilots Fail (And What to Automate First)
Author
Grovant Editorial · AI Practice
Published
Aug 17, 2026
Reading time
8 min read

Let's first state the obvious: everyone is doing something with AI right now.

McKinsey's State of AI survey has organizational AI adoption at 78 percent, meaning nearly four out of five companies report using AI in at least one business function. Your competitors are in that number. So are you, probably, even if it is just people quietly pasting things into a chatbot.

Now here is the number nobody puts on the conference slide. Researchers at MIT studied hundreds of corporate generative AI deployments and found that about 95 percent of pilots produced no measurable impact on profit and loss. Ninety-five percent. Adoption is everywhere. Returns are almost nowhere.

So what separates the 5 percent that works from the 95 percent that quietly dies in a slide deck?

We build AI automations for businesses, which means we spend a lot of time cleaning up after failed pilots. The pattern behind the failures is remarkably consistent, and it has almost nothing to do with the technology. In this article, we'll break down why pilots fail, what Klarna's very public AI story teaches, and exactly how to pick your first automation. Let's dive in.

By the numbers

95%

AI pilots with no measurable P&L impact

MIT research, reported August 2025

78%

Organizations using AI in at least one function

McKinsey State of AI

~280x

Drop in inference cost, Nov 2022 to Oct 2024

Stanford AI Index 2025, GPT-3.5-level performance

Why pilots fail: they start with the technology

Here's how the typical failed pilot is born. Someone senior sees a demo. The demo is genuinely impressive, because demos always are. A team gets assigned to "do something with AI," a tool gets bought, and three months later there is a chatbot nobody uses sitting on the intranet.

Notice what never happened in that story: nobody mapped a workflow. Nobody asked which specific, repeated task eats the most hours, what it costs today, and what a fixed version would be worth.

The MIT researchers found the same thing. The failures were not model failures. Companies stalled on integration, on workflows that did not fit how work actually happens, and on tools that never learned from feedback. The 5 percent that succeeded picked one specific process, wired the AI into it properly, and measured before and after.

The technology was the same for both groups. The order of operations was not.

What Klarna's AI story actually teaches

Klarna is the most useful AI case study in business right now, and most people only know half of it.

The half everyone knows: in early 2024, Klarna announced its AI assistant was handling two thirds of all customer service chats in its first month, doing the work of roughly 700 full-time agents, with resolution times dropping from 11 minutes to under 2.

The half fewer people know: about a year later, Klarna's CEO publicly walked part of it back and began recruiting human agents again, saying cost had been over-weighted as a factor and that customers should always have the option of a person.

It would be easy to read that as "AI failed." It did not. The automation still handles a massive share of routine volume. What failed was the framing: automate the job instead of automate the workflow.

The lesson is not "don't automate." The lesson is that the durable wins came from automating the repetitive two thirds and keeping humans on the third that needs judgment. That split, not the headline, is the actual playbook.

What to automate first

Forget the moonshots for a minute. Your first automation should be boring. Boring is where the return is.

Three questions pick the winner. How often does the task happen? How rule-based is it? And what does a mistake cost?

You want high frequency, mostly rules, and cheap mistakes. A task that happens 40 times a week, follows the same steps every time, and gets reviewed by a human anyway is a perfect first target. A task that happens twice a quarter and torches a client relationship if it goes wrong is a terrible one, no matter how impressive the demo looks.

WorkflowWhy it works as a first automation
Lead intake and routingHigh volume, clear rules, a human still makes the call
Invoice and document processingRepetitive extraction, verifiable against the source
First-draft reportingAI drafts from your data, a person approves and sends
Support triage and taggingRoutes and prioritizes, humans handle the judgment tier
Data entry between toolsPure repetition, the current process is copy-paste
Good first automations share a shape: frequent, rule-based, human-reviewed.

Notice what is not on that list. Anything client-facing that requires judgment. Anything where the AI's output ships without review. Anything strategic. Those can come later, after the boring wins have paid for the learning curve and built trust with your team.

The economics quietly changed in your favor

There is one more reason the timing matters. The cost side of this equation has collapsed.

According to Stanford's 2025 AI Index, the cost of running a model at GPT-3.5 level performance fell roughly 280-fold between November 2022 and October 2024. Automations that were economically silly two years ago now cost pennies per run.

This is why we quote automations in cost per run, not vague monthly platform fees. When an invoice-processing workflow costs a fraction of a cent in tokens per document, you can put an exact number on what the automation costs against what the manual hour costs. The math stops being a leap of faith and becomes arithmetic.

How to run a pilot that actually ships

The 5 percent that succeed follow roughly the same sequence, and none of it is glamorous.

  1. Map the workflow first. Write down every step of the process as it happens today, who does it, and how long it takes. If you cannot map it, you cannot automate it.
  2. Measure the baseline. Hours per week, error rate, turnaround time. Without a before number, your pilot cannot prove anything, and unprovable pilots get cancelled.
  3. Pick one workflow, not a platform. A pilot that automates one process end to end beats a platform rollout that automates nothing in particular.
  4. Keep a human in the loop where judgment lives. The Klarna split: automate the repetitive share, route the rest to people.
  5. Ship to production, not to a demo. A pilot that never touches real work produces no evidence. Small scope in production beats big scope in a sandbox.
  6. Compare against the baseline after 30 days, then decide: expand, adjust, or kill it. All three are wins if the numbers are real.

Start with the map, not the model

Everything above compresses into one sentence: companies that map workflows and measure baselines end up in the 5 percent, and companies that buy tools and hope end up in the 95.

The map is where we start every engagement, and it is the part we give away. A senior automation engineer walks through how your team actually works, finds the repetitive workflows worth automating, and hands you a ranked list with the estimated cost per run for each. It is free, it takes a few days, and the map is yours whether or not you build any of it with us.

Because if you win, we win. And the fastest way for both of us to lose is to start with a demo instead of a map.

Signed
Grovant Editorial · AI Practice
Filed in Operations · 8 min read
Back to the Journal
Done-for-you

Want us to deliver this for you?

One senior specialist owns the work start to finish. Tell us what you need and get a free, no-pressure proposal.

What do you need?

No spam. A senior partner replies within 1 business day.

Practice intake

Hand us the brief. Senior reply in one business day.

Skip the SDR loop. The senior who would run your engagement reads every submission and writes back inside one business day.

  • Senior practice lead replies, not a coordinator
  • Mutual NDA on request, signed before email two
  • Free 4-business-day audit, no follow-up sequence

Send it now

Read by a senior · one business day · no SDR loop

Article FAQ

Frequently asked questions.

Quick answers to what readers ask about this topic.

  • MIT research covering hundreds of corporate deployments found roughly 95 percent of generative AI pilots produced no measurable P&L impact. The failures were rarely about model quality. Companies stalled on integration, picked workflows that did not match how work actually happens, and never measured a baseline, so the pilot could not prove anything either way.

  • Start with tasks that are high frequency, rule-based, and cheap to get wrong: lead intake and routing, invoice and document processing, first-draft reporting, support triage, and data entry between tools. Keep a human reviewing the output. Client-facing judgment calls and anything that ships without review should come later.

  • Klarna's AI assistant handled two thirds of customer service chats in its first month, equivalent to about 700 agents' workload. Roughly a year later the company started hiring human agents again, saying cost had been over-weighted and customers should always be able to reach a person. The automation still handles routine volume; the correction was about keeping humans where judgment matters.

  • Far less than most people assume. Stanford's 2025 AI Index found inference costs for GPT-3.5-level performance fell roughly 280-fold between late 2022 and late 2024. Well-built automations typically cost fractions of a cent to a few cents per run in model usage, which is why we quote cost per run rather than opaque platform fees.

  • One workflow, in production, measured against a baseline within about 30 days. If a pilot has run longer than a quarter without touching real work or producing a before-and-after number, it is not a pilot anymore. Small scope in production beats big scope in a sandbox.

  • A senior automation engineer walks through how your team actually works, documents the repetitive workflows, and ranks which ones are worth automating with an estimated cost per run for each. Grovant runs these free, and the map is yours to keep whether or not you build the automations with us.

Don't see your question?Send a quick message →
Reply · within 1 business day

Want to talk through this with a senior owner? Send the brief.

Send the brief