Let's first state the obvious: everyone is doing something with AI right now.
McKinsey's State of AI survey has organizational AI adoption at 78 percent, meaning nearly four out of five companies report using AI in at least one business function. Your competitors are in that number. So are you, probably, even if it is just people quietly pasting things into a chatbot.
Now here is the number nobody puts on the conference slide. Researchers at MIT studied hundreds of corporate generative AI deployments and found that about 95 percent of pilots produced no measurable impact on profit and loss. Ninety-five percent. Adoption is everywhere. Returns are almost nowhere.
So what separates the 5 percent that works from the 95 percent that quietly dies in a slide deck?
We build AI automations for businesses, which means we spend a lot of time cleaning up after failed pilots. The pattern behind the failures is remarkably consistent, and it has almost nothing to do with the technology. In this article, we'll break down why pilots fail, what Klarna's very public AI story teaches, and exactly how to pick your first automation. Let's dive in.
95%
AI pilots with no measurable P&L impact
MIT research, reported August 2025
78%
Organizations using AI in at least one function
McKinsey State of AI
~280x
Drop in inference cost, Nov 2022 to Oct 2024
Stanford AI Index 2025, GPT-3.5-level performance
Why pilots fail: they start with the technology
Here's how the typical failed pilot is born. Someone senior sees a demo. The demo is genuinely impressive, because demos always are. A team gets assigned to "do something with AI," a tool gets bought, and three months later there is a chatbot nobody uses sitting on the intranet.
Notice what never happened in that story: nobody mapped a workflow. Nobody asked which specific, repeated task eats the most hours, what it costs today, and what a fixed version would be worth.
The MIT researchers found the same thing. The failures were not model failures. Companies stalled on integration, on workflows that did not fit how work actually happens, and on tools that never learned from feedback. The 5 percent that succeeded picked one specific process, wired the AI into it properly, and measured before and after.
The technology was the same for both groups. The order of operations was not.
What Klarna's AI story actually teaches
Klarna is the most useful AI case study in business right now, and most people only know half of it.
The half everyone knows: in early 2024, Klarna announced its AI assistant was handling two thirds of all customer service chats in its first month, doing the work of roughly 700 full-time agents, with resolution times dropping from 11 minutes to under 2.
The half fewer people know: about a year later, Klarna's CEO publicly walked part of it back and began recruiting human agents again, saying cost had been over-weighted as a factor and that customers should always have the option of a person.
It would be easy to read that as "AI failed." It did not. The automation still handles a massive share of routine volume. What failed was the framing: automate the job instead of automate the workflow.
The lesson is not "don't automate." The lesson is that the durable wins came from automating the repetitive two thirds and keeping humans on the third that needs judgment. That split, not the headline, is the actual playbook.
What to automate first
Forget the moonshots for a minute. Your first automation should be boring. Boring is where the return is.
Three questions pick the winner. How often does the task happen? How rule-based is it? And what does a mistake cost?
You want high frequency, mostly rules, and cheap mistakes. A task that happens 40 times a week, follows the same steps every time, and gets reviewed by a human anyway is a perfect first target. A task that happens twice a quarter and torches a client relationship if it goes wrong is a terrible one, no matter how impressive the demo looks.
| Workflow | Why it works as a first automation |
|---|---|
| Lead intake and routing | High volume, clear rules, a human still makes the call |
| Invoice and document processing | Repetitive extraction, verifiable against the source |
| First-draft reporting | AI drafts from your data, a person approves and sends |
| Support triage and tagging | Routes and prioritizes, humans handle the judgment tier |
| Data entry between tools | Pure repetition, the current process is copy-paste |
Notice what is not on that list. Anything client-facing that requires judgment. Anything where the AI's output ships without review. Anything strategic. Those can come later, after the boring wins have paid for the learning curve and built trust with your team.
The economics quietly changed in your favor
There is one more reason the timing matters. The cost side of this equation has collapsed.
According to Stanford's 2025 AI Index, the cost of running a model at GPT-3.5 level performance fell roughly 280-fold between November 2022 and October 2024. Automations that were economically silly two years ago now cost pennies per run.
This is why we quote automations in cost per run, not vague monthly platform fees. When an invoice-processing workflow costs a fraction of a cent in tokens per document, you can put an exact number on what the automation costs against what the manual hour costs. The math stops being a leap of faith and becomes arithmetic.
How to run a pilot that actually ships
The 5 percent that succeed follow roughly the same sequence, and none of it is glamorous.
- Map the workflow first. Write down every step of the process as it happens today, who does it, and how long it takes. If you cannot map it, you cannot automate it.
- Measure the baseline. Hours per week, error rate, turnaround time. Without a before number, your pilot cannot prove anything, and unprovable pilots get cancelled.
- Pick one workflow, not a platform. A pilot that automates one process end to end beats a platform rollout that automates nothing in particular.
- Keep a human in the loop where judgment lives. The Klarna split: automate the repetitive share, route the rest to people.
- Ship to production, not to a demo. A pilot that never touches real work produces no evidence. Small scope in production beats big scope in a sandbox.
- Compare against the baseline after 30 days, then decide: expand, adjust, or kill it. All three are wins if the numbers are real.
Start with the map, not the model
Everything above compresses into one sentence: companies that map workflows and measure baselines end up in the 5 percent, and companies that buy tools and hope end up in the 95.
The map is where we start every engagement, and it is the part we give away. A senior automation engineer walks through how your team actually works, finds the repetitive workflows worth automating, and hands you a ranked list with the estimated cost per run for each. It is free, it takes a few days, and the map is yours whether or not you build any of it with us.
Because if you win, we win. And the fastest way for both of us to lose is to start with a demo instead of a map.