Do 95% of AI pilots really fail?
Nobody knows the real rate. The "95% of AI pilots fail" line comes from MIT NANDA's 2025 report, and it rests on a review of 300+ public AI projects, interviews at 52 organizations and 153 surveys collected at conferences.
That's a small sample for a number that ends up in so many board decks, and Wharton's Kevin Werbach challenged it publicly. The same report has a finding that gets quoted less. Projects bought from outside partners or built with them succeeded about twice as often as internal builds, about 67% against 33%. Those figures are self-reported too. We are an outside partner, so give that number the same discount you give the 95%.
In RSM's July 2026 survey of 1,030 middle-market companies in the US and Canada, the top blockers among those with only moderate or limited pilot success were data quality (53%) and integration (47%). Across all 1,030 companies, 58% plan to spend $1M or more on AI in 2026, and 52% already use outside consultants.
Why AI pilots fail before they reach production
Ask why AI pilots fail and people talk about the model. Spend time with a stalled pilot and you find the problem somewhere around it.
Part of the problem is the calendar. Mid-market PE funds expect AI proofs of concept in 2-4 weeks, according to Korn Ferry (September 2026). A proof of concept in that window is realistic. Calling it production is where the trouble starts.
It ran on demo data. The pilot was built and tested on a clean export or a folder of hand-picked examples. Real inputs are phone photos, spreadsheets with merged cells, forwarded threads with the attachment buried several replies down, and PDFs with another document stapled to the end. A pilot that never saw the mess has nothing useful to say about it.
The documented process isn't the real one. Every team has exceptions that live in someone's head: the customer who always sends the wrong form, the approver who has to see anything from a particular supplier. A pilot built from the procedure document misses all of them. The only fix we know is to sit with the people doing the work and watch.
Nobody owns it. The pilot belonged to a central AI team or to IT, and the operations lead whose numbers it would change never signed up. When it stalls, nobody's week gets worse.
There was no baseline. If nobody measured how long the work takes today and how often it goes wrong, "it works" is an opinion. Opinions lose budget meetings.
No evals. An eval set is a collection of real cases with the correct answer written down, scored every time a prompt or a model changes. Without an eval set, each change is a guess, and a single bad answer in front of the CFO can end the project.
It never got into the system of record. The pilot read a spreadsheet. Production has to read from your ERP or CRM and write back to it, with permissions someone has to approve. That's the integration blocker from the RSM survey, and access requests eat calendar time, so start them on day one.
Security saw it last. When the pilot was "ready", the security review found client data going to a model provider nobody had approved, and the project went back to the start. A review at the beginning costs a meeting. At the end it costs the project.
The handover was an afterthought. An engineer working alone or an outside vendor built it, and nobody else can change it. Gartner expects that by 2028, 70% of enterprises will abandon agentic AI built by vendor forward deployed engineers, because costs climb and the companies can't evolve the systems themselves (The Register, 30 September 2026).
How to implement AI in your business: a pilot-to-production checklist
Run this list before you restart a stalled pilot or approve a new pilot. Every line you can't tick is a place the next pilot can stall.
- One named workflow with a clear start and end. "Invoice arrives" to "bill approved in NetSuite" is a workflow. "AI for finance" is not.
- A business owner whose numbers move if it works, and who agreed to the target.
- Baseline numbers from today: volume, time per item, error rate and cost per item.
- The real process, mapped by watching the work, with the exceptions written down.
- A call on every step: delete it, plain code, AI, or a human decision. Our accounts payable breakdown shows what that looks like across a full workflow.
- Access to the system of record agreed early, through a service account limited to what the workflow needs.
- An eval set built from real cases, with a pass mark the owner signed off before the build started.
- A security and data review before the build, covering which data goes to which model provider and where the logs live.
- Human approval wherever money or risk is involved.
- An audit trail showing what the AI saw and proposed, and who approved it.
- Shadow mode on live work, with outputs compared against what your team did, before anything goes live.
- A written handover, and a named person on your side who can change the workflow without us.
- An exit plan: who owns the code and the eval set, and how you would run it if we stopped working together.
Gartner's advice to buyers of vendor-built AI covers the same ground: settle scope, incentives, governance, ownership, knowledge transfer and the exit plan from day one.
Knowledge transfer is the easiest line to skip, and most buyers do skip it: in BCG's January 2026 survey of PE investors, only 45% make sure the knowledge reaches their own teams.
How long does it take to move a stalled pilot into production?
With us it takes a 2-3 week Deployment Diagnostic, then a 6-8 week Production Sprint, both at fixed prices. Part of the Diagnostic's job is to tell you whether the pilot deserves saving.
The way we run it has a name now: forward-deployed engineering. A senior engineer works inside your team and your systems, with the studio behind them.
| Stage | What happens | At the end you have |
|---|---|---|
| Deployment Diagnostic (2-3 weeks, $7,500 fixed) | A senior engineer interviews the people doing the work and compares the documented process with the real one, inside your actual systems. They set the baseline numbers and decide where AI belongs. | A process map, baseline numbers, an eval plan and a fixed quote. The $7,500 is credited against the Sprint if you go ahead. |
| Sprint: build | We build inside the tools your team already uses, on real data, with the eval set scoring every version. | A working version with an eval score the owner can read. |
| Sprint: controls | We add the audit trail and human approval wherever money or risk is involved, and we limit the system account to what the workflow needs. | Controls your auditors and security team have seen before anything goes live. |
| Sprint: shadow mode | The workflow runs on live work beside your team, and we compare its output with theirs, case by case. | Numbers against the baseline, and a go or no-go call from the business owner. |
| Go-live and handover | It goes live when the owner signs off on the shadow-mode numbers. We write the handover for whoever maintains it next. | A live workflow and a written handover. If you want us to keep it running, the AI Team retainer starts from $3,000 per month. |
The Production Sprint is $25,000-$60,000 fixed, quoted at the end of the Diagnostic. Where it lands in that range depends on how many systems it touches and how many input formats it has to read.
From your side, the Sprint needs a business owner who can make decisions quickly, access through your IT team, a security contact who signs off on data handling, and regular time with the people who do the work. When one of those goes missing, the Sprint slips, and we say so as soon as we see it.
For context, small AI studios quote a paid diagnostic at around $5,000-$15,000 and one workflow to production at around $25,000-$50,000. Gartner puts consulting fees for vendor forward deployed engineers at up to about $200,000 per quarter per use case (Channel Dive, 30 September 2026).
If you'd rather keep someone inside the team for longer, our Embedded Engineer option puts a senior DK engineer with you 2-3 days a week for $16,000-$20,000 per month.
When to kill an AI pilot instead
Kill it when the Diagnostic shows the work isn't worth automating or nobody will own the result. A dead pilot that frees budget for the right workflow is a good outcome. The money already spent on the pilot is gone either way. What matters is whether the next dollar gets a workflow into production.
- The step shouldn't exist. Automating a report nobody reads gets you a faster report nobody reads.
- Plain code does the job. If the rule fits in an if-statement, a model adds cost and a new way to be wrong.
- The volume is too low. A task that comes up rarely won't pay back a build. Write a checklist for it instead.
- Nobody will own it. Without a business owner it will stall again, whoever builds it.
- The data doesn't exist. If the information the AI needs isn't recorded anywhere, your first project is recording it.
- Checking costs as much as doing. If a person has to redo every answer to trust it, you've added a step.
- You already pay for it. Check what your ERP or helpdesk vendor released recently before you build a copy.
We'd rather tell you to kill a pilot at the end of a Diagnostic than sell you a Sprint that stalls halfway. Once you've picked a first workflow, our AI implementation services start with that triage. Still choosing? That is what the five-day AI Readiness Audit, from $800, is for.
Accounts payable is a good first candidate, because the volume is high and the human checkpoint is obvious: payment approval. We sorted each AP step in AI accounts payable automation. Professional-services firms can see the same approach applied to client work in AI for accounting firms.