Economics and buying Playbook

AI Pilot to Production: How to Rescue, Restart, Buy or Stop a Stalled Pilot

In the largest surveys about half of AI proofs of concept reach production, and about one organization in twenty gets material financial value from AI. The gap between those numbers is made of decisions a CTO controls: an owned metric, production data, written gates and costs modeled at volume.

For CTOs, CIOs and heads of AI with a pilot that worked in the demo and has not been signed off to run for real, deciding whether to rescue it, rebuild it, buy instead or stop.

Published
Reviewed
Reading time
22 min

The short answer

Roughly half of AI proofs of concept reach production in the largest surveys, and about one organization in twenty gets material financial value from AI. Pilots stall on four gaps: no owned business metric, data that is unavailable in production, risk controls added late, and costs never modeled at volume. Close those with written go-live gates, then give each stalled pilot one of four exits: rescue, restart, buy or stop.

Key takeaways

  • In the surveys that disclose their samples, about half of AI proofs of concept reach production: 48% in Gartner's and 46% in IDC's survey for Lenovo.12
  • Value is the rarer event. McKinsey's AI high performers, who attribute at least 5% of EBIT to AI, have held at about 6% of organizations for two years.3
  • The famous "95% of pilots fail" sentence has no derivation in MIT NANDA's own preliminary report, and RAND's "80%" relays a 2022 news estimate.45
  • Token prices keep falling and bills keep rising. Gartner expects inference cost per agentic workflow to grow more than fivefold through 2028, so a business case priced at pilot volumes understates production cost.67
  • Partnering went with roughly twice the deployment rate in MIT's sample, and Gartner predicts 70% of enterprises will abandon agentic AI built by vendor engineers that they cannot evolve themselves. Bring in a partner for the transition and own the system at the end.48

Most companies that started an AI pilot in 2024 or 2025 have one in the same state today. The demo impressed the leadership team, a small group kept it alive, and nobody has signed the decision to run it for real. In a Zapier survey of more than 800 senior leaders, 84% had at least one AI pilot that never reached production, and 38% said their longest-running pilot had been in evaluation for over a year.9 The explanation people reach for is a statistic, usually that 95% of AI pilots fail. That figure is weaker than its fame, and the stronger numbers point somewhere more useful. Reaching production is close to a coin toss. Paying back once there is the hard part, and both outcomes turn on choices made before and during the pilot.

  • 48%of AI projects reached production in Gartner's survey of 644 organizations, about eight months after the prototype1
  • 6%of organizations attribute 5% or more of EBIT to AI, a share that has not moved in two years3
  • 5xGartner's predicted rise in inference cost per agentic workflow through 2028, while token prices keep falling6

How many AI pilots actually reach production

"How many AI pilots fail" hides two questions, and the evidence answers them very differently. The first is whether a pilot becomes a running system. Gartner surveyed 644 organizations in the US, Germany and the UK at the end of 2023 and found that on average 48% of AI projects made it into production, about eight months after the prototype.1 S&P Global's 451 Research surveyed 1,006 professionals in North America and Europe a year later and found the average organization ended 46% of its proofs of concept before production, which leaves a little over half going through.10 IDC's 2026 CIO Playbook for Lenovo, based on 3,120 decision-makers surveyed in late 2025, reports that 46% of AI proofs of concept progressed into production.2 Deloitte sets a stricter bar and gets a lower number: 25% of 3,235 leaders say their organization has moved 40% or more of its AI pilots into production.11

The second question is whether the system pays once it runs, and here the answer is about one in twenty. McKinsey's State of AI 2026 found 44% of respondents reporting AI scaling across the enterprise, while only 37% attribute any EBIT impact to AI and most of those put it under 5%. Its high performers, who attribute 5% or more of EBIT to AI, have stayed at about 6% for two years.312 BCG's 2025 study of 1,250 executives puts 5% of firms in its "future-built" group and 60% getting hardly any material value.13 Gartner's 2026 survey of 1,303 organizations found 22% have scaled AI across multiple business units, and 11% could not say what their function had spent on AI the year before.14

Exhibit 1Reaching production is common; paying back is rare
  • Gartner: AI projects that reached production48%
  • IDC for Lenovo: proofs of concept that progressed to production46%
  • McKinsey: organizations attributing any EBIT impact to AI37%
  • Gartner: organizations that scaled AI across several business units22%
  • McKinsey: high performers, 5% or more of EBIT from AI6%
  • BCG: future-built firms getting value at scale5%
Self-reported surveys, 2023 to 2026. The first two rows count projects and the rest count organizations, which is why the figures answer different questions. Sources: [1], [2], [3], [14], [13]

Read together, the numbers move the bottleneck. Getting from pilot to production succeeds about half the time. Getting from production to a result the finance team recognizes succeeds for a small minority, and that minority has not grown while adoption has. Every figure here is self-reported or an analyst's prediction, and no one independently audits enterprise AI outcomes, so treat them as the best available signal and plan with them on that basis.

What the famous failure numbers measured

Three statistics dominate this topic, and each one fails the simplest test a CTO can apply: what did the source actually count? The table sets each widely quoted figure beside what its publisher measured.

Figure as usually quotedSourceWhat was measuredHow to use it
"95% of GenAI pilots fail"MIT NANDA, July 2025, labeled preliminary findingsNo passage derives the 95%. The nearest figure is a funnel for task-specific tools: 60% of organizations evaluated, 20% piloted, 5% reported sustained impact about six months later. Based on 52 interviews and 153 survey responses4A directional interview study; about one in four piloting organizations cleared its bar
"80% of AI projects fail"RAND, August 2024An estimate relayed from a 2022 news article. RAND's own work is 65 interviews about root causes5Cite RAND for the root causes
"88% of proofs of concept never reach production"IDC for Lenovo, reported March 2025Vendor-sponsored 2024 ratios15Superseded by the same sponsor's 46% in January 20262
"42% of companies abandoned most AI initiatives"S&P Global, 451 ResearchCompanies, up from 17% a year earlier, in a survey of 1,00610Sound, quoted with its definition
"30% abandoned after proof of concept" and "40% of agentic projects canceled"Gartner, 2024 and 2025Predictions with no published sample1617Quote them as predictions
The two figures that hold up best both put production at about half. The MIT figure circulates most and rests on the least.

The MIT report deserves a closer look because it shaped a year of coverage. Its own exhibit carries the authors' caveats: the numbers are directional, they come from individual interviews, and six months may be too short to judge complex systems.4 Ray Poynter pointed out that if 20% piloted and 5% cleared the bar, roughly one piloting organization in four succeeded.18 Wharton's Kevin Werbach said he could not find where the 95% came from.19 The report still has useful findings, including the "learning gap" of tools that never retain feedback, and the observation that workers at more than 90% of firms use personal AI tools while only 40% of firms bought an official subscription.4

A defensible sentence replaces all three headlines. In the largest surveys roughly half of AI proofs of concept reach production, and about 6% of organizations attribute 5% or more of EBIT to AI. Leading a board paper with "95% fail" tells anyone who has read the source that the paper's author has not.

Why pilots stall between the demo and production

The research on causes points at organizations and data far more than at models. RAND's five root causes are a misunderstood problem, missing data, technology chosen before the user's problem, inadequate infrastructure to manage data and deploy models, and problems too hard for current AI; 84% of its interviewees named a leadership-driven cause as the main one.5 Gartner's four reasons for abandonment are poor data quality, inadequate risk controls, escalating costs and unclear business value.16 In a separate Gartner survey, 63% of organizations either lacked the right data management practices for AI or were unsure whether they had them.20 Zapier found infrastructure, data quality and integration problems blocking 41% of stalled pilots and legal or compliance concerns blocking 29%, and organizations that deploy most of their pilots were twice as likely to have an executive sponsor.9

Deloitte describes the structural trap well. A small team can run a pilot in a few months on cleansed data in an isolated environment. Production needs infrastructure, integration with existing systems, security reviews, compliance checks, monitoring and maintenance, and failures that were learning opportunities in the pilot become business risks once real customers depend on the system.21 A pilot built on an extract, with a broad service account and no cost ceiling, proves the model can answer. It says very little about whether the system can run.

Adoption kills pilots more quietly. The shadow-AI finding shows the demand is real, and a sanctioned tool that cannot remember context, learn from corrections or reach the right data loses to the consumer chatbot people already use.4 Licensed assistants show the same pattern: in a 2025 Gartner survey of IT leaders, only 5% of organizations that had finished a Microsoft 365 Copilot pilot were moving to a larger deployment.22

What production asks of a system that a demo never did

A pilot proves a model can give good answers on a curated set. Production has to prove the whole chain of retrieval, prompts, model, tools and gateway holds its quality on live inputs, survives attack, stays inside quota and budget, and keeps working when the provider changes the model underneath it. Microsoft's architecture guidance warns that a team can be mature at MLOps and a beginner at operating generative AI, because the governed system now includes the orchestrator, the calls behind it and the prompts it builds.23 Google's guidance asks for evaluation of whole chains, checks for skew between evaluation inputs and production inputs, and strict versioning of every component that can change.24 Our evaluation guide covers how to build that test suite.

DimensionIn the pilotIn production
EvaluationA curated set of examplesContinuous evaluation on sampled live traffic, with skew checks against the test set24
Unit of changeThe promptThe whole chain: templates, retrieval index, embedding model and model version, each versioned24
ModelFixed for the demoA dependency with a retirement date and occasional silent regressions25
LoadA few usersQuotas, throttling, latency targets and failover behind a gateway23
IdentityOne broad service accountLeast privilege, separate environments, approval for destructive actions26
DataA static extractPipelines that keep indexes fresh and honor deletion requests23
SecurityNever attackedIndirect prompt injection through every email and document the system reads27
ObservabilityConsole logsStep-level traces of every agent run28
Human reviewThe pilot team reads outputsA queue for low-confidence and high-stakes cases, with reviewers' time budgeted29
Nine disciplines a demo never exercises. Each one is cheap to design in and expensive to retrofit.

The model row surprises most teams. On Microsoft Foundry, a generally available model version retires 18 months after launch, every request to it fails from that day, and the official replacement is named only about 90 to 120 days before.25 Providers also regress without warning. Anthropic's September 2025 postmortem described infrastructure bugs that at the worst hour affected 16% of requests to one model, and said its evaluations had not captured what users were reporting.30 OpenAI rolled back a model update within four days after A/B tests, offline evaluations and expert reviews all missed a change in its behavior.31 If the labs' own evaluations miss regressions, a pilot's golden set will miss them too, which is why production evaluation runs on live traffic and never stops.

Agents add durable state, compounding errors and a much larger attack surface. Anthropic's account of its multi-agent research system says the last mile often becomes most of the journey, that one failed step can send an agent down a different path, and that full production tracing was how its team diagnosed failures.28 EchoLeak, a zero-click prompt injection in Microsoft 365 Copilot rated 9.3 for severity, showed how one crafted email could pull internal data out through an assistant; Microsoft fixed it server-side and reported no exploitation.27 OWASP's 2025 list for LLM applications names excessive agency and unbounded consumption among its top risks, and both are invisible in a demo.26 Our guide to agent guardrails sets out the controls.

The bill that only appears at volume

Token prices for a fixed level of capability are collapsing, by 9 to 900 times a year depending on the task in Epoch AI's measurement.7 Bills rise anyway. Gartner calls this the inference paradox: cheaper models tempt teams into agentic designs that reason, retry and check their own work, and it predicts that inference cost per agentic workflow will rise more than fivefold through 2028.6 It also expects inference to make up at least 70% of a model's lifetime cost, and at least half of generative AI projects to overrun their budgets through 2028 because of poor architectural choices and thin operating experience.32

Exhibit 2Tokens used per task, compared with a chat exchange
  • Chat exchange1x
  • Single agentabout 4x
  • Multi-agent systemabout 15x
Anthropic's measurement from its own research system. Multi-agent designs pay off only where the task is worth the tokens. Source: [28]

The clearest public evidence came from internal developer tools. Uber opened agentic coding tools to most of its engineers in December 2025, reportedly used up its 2026 AI budget by April, and then capped spending per tool.33 On the customer side, Gartner predicts that by 2030 the generative AI cost per customer-service resolution will exceed what many offshore human agents cost, as vendors move from subsidized growth to profit.34 A pilot business case built on list token prices and chat-length tasks understates production cost three ways at once: fifty pilot users become five thousand, agent loops multiply the tokens per task, and vendor pricing is heading up.

One test ties the cost drivers together: the cost per successful task at projected volume, set against the measured cost per task today, counting staff time, system cost, rework and the expected loss from errors. Gartner warned in 2024 that CIOs who do not understand how their generative AI costs scale could miscalculate them by 500% to 1,000%, so the pilot is the place to measure both sides.35 Our guides to AI ROI and inference cost show how to build both sides of that comparison.

Public failures and the control each one lacked

There are now enough public cases of AI systems withdrawn, rolled back or ruled against to show a pattern. In almost every one, the model's raw capability was a minor factor. A control was missing.

CaseWhat happenedThe control that would have caught it
Air Canada, 2024Its website chatbot told a customer bereavement fares could be claimed after travel, against the airline's own policy. The tribunal rejected the argument that the chatbot was responsible for its own words36Answers grounded in the authoritative policy, with review of anything that commits the company
McDonald's and IBM, 2024A voice-ordering test in more than 100 US restaurants ended in July 202437An accuracy bar measured on live orders before expanding
New York City MyCity, 2024 to 2026The city's business chatbot told owners they could take workers' tips; disclaimers were added, and the bot was shut down in 20263839Grounding and review for legal answers, and written stop criteria
Klarna, 2024 to 2025Its assistant handled two-thirds of service chats in its first month; in 2025 Klarna restored a human option after cost-led design lowered quality4041A quality metric beside the cost metric, and a route to a person
Commonwealth Bank of Australia, 2025Cut 45 roles after launching a voice bot, then reversed the decision as call volumes rose42A value case measured on operating data before staffing decisions
Replit, 2025A coding agent deleted a live database during a declared code freeze; the company then separated development and production automatically43Least privilege, environment separation and enforced approval
Cursor, 2025A support bot invented a one-device policy, and customers canceled44Grounding, disclosure and review of statements that affect accounts
Deloitte Australia, 2025A government review contained fabricated references; Deloitte disclosed its AI toolchain and refunded part of the fee45A verification step for every citation
Starbucks, 2026An AI inventory-counting tool that scored 99% in controlled tests was scrapped nine months after launch, as reflections doubled milk counts and syrups were misidentified on real shelves46Accuracy measured in real store conditions, and a retraining plan for changing stock
Public reporting and the Air Canada tribunal decision. Most of these ended in redesign, and the missing controls are cheap next to the exposure.

Two details matter more than the headlines. Most of these cases ended in redesign: Klarna runs AI and people side by side, and Commonwealth Bank kept its bot and reversed only the staffing decision.4142 And a popular retelling is wrong. Klarna's "700 agents" was a workload equivalent in its 2024 announcement, and its filing with the SEC reports the assistant handling 66% of service chats in the year to March 2025.4047 Quoting these cases accurately is part of being taken seriously by a board that has read them too.

What the organizations that scale do differently

Workflow redesign is the most consistent difference. In McKinsey's 2026 survey nearly three-quarters of high performers had fundamentally redesigned workflows around AI, against about a quarter of everyone else.3 Its March 2025 edition tested twelve adoption practices and found that tracking well-defined KPIs had the largest effect on the bottom line, while fewer than one organization in five did it.48 High performers are also more likely to have defined rules for when a model's output needs human validation.12

Exhibit 3Where the firms that get value from AI differ
  • McKinsey high performers that redesigned workflowsnearly 3 in 4
  • Other organizations that redesigned workflowsabout 1 in 4
  • BCG future-built firms: initiatives deployed62%
  • BCG laggards: initiatives deployed12%
  • BCG future-built firms that rigorously track AI value60%+
  • BCG stagnating firms that rigorously track AI value17%
Self-reported surveys from 2025 and 2026. They show strong association, and no study yet proves these practices cause the results. Sources: [3], [13]

BCG adds the organizational detail. Future-built companies are 1.5 times as likely to share AI ownership between business and IT, and BCG names exclusive IT ownership as a marker of stagnating firms. They also concentrate: 62% of their initiatives are deployed, against 12% at laggards that spread effort across scores of workflows.13 MIT's successful buyers found use cases through frontline managers, judged tools on operating outcomes instead of model benchmarks, and started at the edge of a workflow before moving inward.4

Uber's code-review system shows the method in practice. It rolled out one team at a time, tracked precision on dashboards, counted false positives reported by engineers, and compared itself with human reviewers. It now covers more than 90% of about 65,000 weekly code changes, and engineers rate 75% of its comments useful.49 The figures are Uber's own, and the method is the lesson: one workflow, one usefulness metric, and expansion only when the number held.

Ten gates a pilot passes before go-live

No single framework publishes a production gate for generative AI, but the public guidance agrees on what one should contain. AWS's Generative AI Lens covers output quality, excessive agency, observability, versioning and inference cost.50 Microsoft's Responsible AI Standard requires release criteria for each metric and error type before release, named owners who can override or stop the system, and a documented fallback for each predictable failure with the time it takes to invoke.51 NIST's Generative AI Profile sets out the measurement and management duties behind them.52 The table combines those sources into the gates we would hold any release to.

GateShip only when
1. Problem and baselineThe problem is written down and today's cycle time, error rate and cost per case are measured
2. Release thresholdsEach metric and error type has a threshold, met on production-like data, with a re-evaluation schedule
3. SecurityRed-team testing has covered indirect prompt injection and excessive agency
4. Unit economicsCost per task and latency stay inside budget at projected volume
5. Versioning and model lifecyclePrompts, chains, index, embeddings and model versions are versioned, and a plan exists for the model's retirement
6. MonitoringRuns are traced end to end, and live traffic is sampled into continuous evaluation with drift checks
7. OwnershipA named person can approve, override and shut down, and the incident runbook has been rehearsed
8. Fallback and rollbackThe fallback is documented with its time to invoke, and the gateway can roll back a release
9. Progressive exposureThe system has run in shadow mode, then on a canary slice of traffic
10. Escalation and autonomyRetries and actions are capped, and people approve irreversible or high-value actions
Assembled from the AWS Generative AI Lens, Microsoft's Responsible AI Standard and Azure guidance, Google Cloud's operating guidance, NIST AI 600-1 and OpenAI's agent guide.505153245229

No vendor-neutral standard sets the numbers, such as a minimum share of answers grounded in a source. The person who owns the use case has to write them, and gate 2 is the one most pilots skip. A pilot without written thresholds can neither pass nor fail, which is exactly how a pilot ends up in evaluation for over a year.

Run the pilot the way production will run

The practical answer is a pilot that already looks like production, run in an order where each step earns the next. Accenture research published in Harvard Business Review found that firms that design for scale from the start attempt scaling, and succeed, about twice as often as firms stuck in a proof-of-concept habit.54

  1. Choose one workflow with an owner Pick a narrow, high-value workflow whose business owner holds the metric, and measure today's baseline before anything is built.
  2. Write the bar first Agree the success metric and the go or no-go threshold, and sign them before the first line of code.
  3. Build on production data Connect through the real integrations and permissions. A cleansed extract proves nothing about the system that will run.
  4. Run in shadow mode The new system receives a copy of live requests while only the current process answers users, so quality is measured on real traffic at no risk.55
  5. Move to approval mode A person approves each action, and every correction becomes a new evaluation case.
  6. Release behind a canary Send a small slice of traffic with rollback ready, widen autonomy only for low-risk actions, and keep irreversible actions with people until reliability is proven.5329

Autonomy is a design decision, separate from what the model can do. A 2025 paper by researchers at the University of Washington describes five levels for the user's role, from operator through collaborator, consultant and approver to observer, and the right level for each action can be set deliberately and raised with evidence.56

Rescue, restart, buy or stop

A stalled pilot has four honest exits. Carrying on as before without new evidence is how pilots reach their second year, so it belongs on no one's list of options.

ExitChoose it whenFirst move
RescueAn owner and a measured baseline exist, production data is reachable, and the gaps are engineering gaps such as evaluation, monitoring, integration, cost routing or governanceAdd the missing gates to the existing build
RestartThe problem is right and the build is wrong: hand-made extracts, sprawling prompts, a stack nobody can operate, or cost and latency far off targetKeep the problem definition, evaluation set and baseline, and rebuild on production data
BuyThe workflow is generic, a mature product exists, and a custom build would not set you apartTest vendors against your evaluation set and agree data portability and exit terms before signing
StopNo owner, no measurable baseline, data unavailable in production, cost per task above value per task at volume, or accuracy stuck below the bar as each fix exposes new failuresRecord what was learned and keep the evaluation set and the data work
Decision signals drawn from RAND, Deloitte, Zapier and practitioner evaluation work.521957

Stopping costs less politically than most teams fear: 81% of Zapier's respondents said a failed pilot had minimal to moderate impact on future AI investment.9 The warning sign for a stop is a plateau. Hamel Husain, who has worked on many production LLM systems, describes the pattern where fixing one failure mode exposes others, and puts error analysis and evaluation at 60 to 80% of development time in the projects he describes.57 Weak evaluation is the most common reason a pilot can neither ship nor be killed.

The evidence on who should build is mixed and worth reading carefully. In MIT's interview sample, external partnerships using customized tools reached deployment about 67% of the time, against about 33% for internal builds, with the authors warning that the figures are self-reported and the correlation may not be causal.4 Menlo Ventures found the share of AI use cases bought instead of built rose from 53% in 2024 to 76% in 2025, and a16z's CIOs said internal tools are hard to maintain as models change; both firms invest in AI application companies.5859 Pulling the other way, 32% of McKinsey's 2026 respondents decided against buying software because they could build it with coding agents.3 And Gartner predicts that by 2028 70% of enterprises will abandon agentic AI built by a vendor's forward-deployed engineers, because of rising costs and an inability to evolve it themselves.8

Our reading of that evidence is simple. Buy where the workflow is generic. Bring in a partner for the production transition, and write the handover of evaluation suites, runbooks and code ownership into the contract. Build where the workflow sets you apart and your team can own evaluation and operations. Our questions for an AI development company cover how to test a partner on exactly that.

What a pilot-to-production assessment should answer

An outside assessment is worth paying for only if it ends in one of the four exits with reasons. It needs people who have run production systems: an AI engineer for prompts, agents, retrieval and tools, a data engineer for pipelines and index freshness, a platform engineer for deployment, cost routing, monitoring and model migration, and a reviewer for security and compliance. Your side keeps the business owner, a domain expert who maintains the evaluation set, and the sponsor. That split matches the skills market: ManpowerGroup's 2026 survey of 39,000 employers ranks AI model and application development as the hardest skill to hire.60

Questions the assessment should settle

  • Who owns the business metric, what is today's baseline, and who sponsors the change?
  • How does the system score on real and edge-case inputs, and what does the error analysis show?
  • Is the data the system needs available in production, fresh, permissioned and governed?
  • What does a successful task cost at projected volume, against what the task costs today?
  • Which actions can the system take, which need a person, and how is it attacked through its inputs?
  • How is it monitored, who is paged, and what happens when the provider retires the model?
  • Rescue, restart, buy or stop, and on what conditions?

The pilot that impressed the board did its job. It showed the idea can work. Whether it becomes a system the business runs on is decided by the next set of choices: a metric someone owns, data the system can reach in production, gates written before the release, and costs measured at the volume the business will actually see.

This is the work behind our AI development service. We take a pilot through the assessment above, write the gates with the person who owns the metric, run shadow and canary releases on production data, and hand over the evaluation suite, runbooks and code so your team owns the system when we step back.

Questions leaders ask

What percentage of AI pilots make it to production?

About half, in the surveys that publish their samples. Gartner found 48% of AI projects reached production in its survey of 644 organizations, and IDC's 2026 survey for Lenovo found 46% of proofs of concept progressed to production. Financial value is rarer: about 6% of organizations in McKinsey's 2026 survey attribute 5% or more of EBIT to AI.

Is it true that 95% of AI pilots fail?

That figure comes from MIT NANDA's July 2025 preliminary report, which never shows how the 95% was derived. Its nearest number is a funnel in which 20% of organizations piloted task-specific tools and 5% reported sustained impact about six months later, so roughly one piloting organization in four cleared its bar. The authors themselves call the figures directional.

Why do AI pilots fail to scale?

Mostly for organizational and data reasons. Research from RAND and Gartner points to unclear business value, missing or poor-quality data, risk controls added late and costs that grow at volume. Pilots built on cleansed extracts with broad permissions also skip the engineering production needs, including live evaluation, versioning, monitoring, security testing and a plan for model retirement.

How long does it take to move an AI pilot into production?

Published figures vary by what they measure. Gartner reported about eight months from prototype to production, BCG reports nine to eighteen months to impact depending on maturity, and Deloitte notes that three-month estimates can stretch to eighteen once integration gets complicated. The honest answer for any one pilot comes from an assessment of its data, integrations and gates, scoped phase by phase.

Should we fix a stalled AI pilot or rebuild it?

Fix it when an owner and baseline exist, production data is reachable and the gaps are engineering ones such as evaluation, monitoring and integration. Rebuild when the problem is right but the build depends on hand-made extracts, brittle prompts or a stack nobody can operate. In both cases keep the problem definition, the evaluation set and the baseline.

Is it better to build, buy or partner for enterprise AI?

Buy where the workflow is generic and a mature product exists. Partner for the move from pilot to production, with the handover of evaluation suites, runbooks and code written into the contract, since Gartner warns that enterprises abandon vendor-built agents they cannot evolve. Build where the workflow sets you apart and your team can own its evaluation and operations.

What should an AI pilot-to-production assessment include?

The business metric, baseline, owner and sponsor; accuracy on real and edge-case inputs with error analysis; production data availability and permissions; cost per task at projected volume against today's cost; security, human review and irreversible actions; monitoring and model retirement plans. It should end with a recommendation to rescue, restart, buy or stop.

Sources

  1. Gartner Survey Finds Generative AI Is Now the Most Frequently Deployed AI Solution in OrganizationsGartner, May 7, 2024
  2. Research From Lenovo Reveals AI Is Paying Off, Yet Most CIOs Aren't Ready for What Comes NextLenovo, January 2026, on IDC's CIO Playbook 2026
  3. The state of AI in 2026: On the road to ROIMcKinsey & Company, August 2026
  4. The GenAI Divide: State of AI in Business 2025MIT NANDA, July 2025 (preliminary findings, copy hosted by AI News)
  5. The Root Causes of Failure for Artificial Intelligence Projects and How They Can SucceedRAND Corporation, August 2024
  6. Agentic AI costs set to balloon fivefold by 2028The Register, August 17, 2026, reporting Gartner
  7. LLM inference prices have fallen rapidly but unequally across tasksEpoch AI
  8. Gartner Predicts 70% of Enterprises Will Abandon Agentic AI Built by Vendor Forward-Deployed Engineering by 2028Gartner, September 29, 2026
  9. 84% of Companies Have Stalled AI Pilots: Here's WhyZapier
  10. Generative AI shows rapid growth but yields mixed resultsS&P Global Market Intelligence, October 2025
  11. From Ambition to Activation: Organizations Stand at the Untapped Edge of AI's Potential, Reveals Deloitte SurveyDeloitte, January 2026
  12. The state of AI in 2025: Agents, innovation, and transformationMcKinsey & Company, November 2025
  13. The Widening AI Value GapBoston Consulting Group, October 2025
  14. Gartner Survey Finds Only 22% of Organizations Have Successfully Scaled AI Across Multiple Business UnitsGartner, September 2026
  15. 88% of AI pilots fail to reach production, but that's not all on ITCIO.com, March 2025
  16. Gartner Predicts 30% of Generative AI Projects Will Be Abandoned After Proof of Concept By End of 2025Gartner, July 29, 2024
  17. Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027Gartner, June 25, 2025
  18. Myth Number 2: MIT Showed That 95% of AI Pilots FailNewMR, Ray Poynter
  19. Why We Don't Believe MIT NANDA's Weird AI StudyFuturiom, August 2025
  20. Lack of AI-Ready Data Puts AI Projects at RiskGartner, February 26, 2025
  21. State of AI in the Enterprise, 2026 editionDeloitte, January 2026
  22. Gartner: Microsoft Copilot hype offset by ROI and readiness realitiestechpartner.news, 2025, reporting Gartner
  23. Generative AI Operations for Organizations with MLOps InvestmentsMicrosoft Learn, Azure Architecture Center
  24. Deploy and operate generative AI applicationsGoogle Cloud Architecture Center
  25. Foundry Models lifecycle and support policyMicrosoft Learn, Microsoft Foundry
  26. OWASP Top 10 for LLM Applications 2025OWASP GenAI Security Project
  27. EchoLeak: The First Real-World Zero-Click Prompt Injection Exploit in a Production LLM SystemarXiv 2509.10540, September 2025
  28. How we built our multi-agent research systemAnthropic Engineering, June 2025
  29. A practical guide to building agentsOpenAI
  30. A postmortem of three recent issuesAnthropic Engineering, September 2025
  31. Expanding on what we missed with sycophancyOpenAI, May 2025
  32. Gartner: Half of Gen AI Projects Could Exceed Budget by 2028Campus Technology, June 22, 2026, reporting Gartner
  33. Uber Caps AI Coding Costs After Exhausting Annual BudgetPYMNTS, June 2026, reporting Bloomberg
  34. Gartner Predicts GenAI Cost Per Resolution for Customer Service Will Exceed Offshore Human Agent Costs by 2030Gartner, January 26, 2026
  35. Gartner Identifies Four Emerging Challenges to Delivering Value from AI Safely and at ScaleGartner, October 21, 2024
  36. Moffatt v. Air Canada, 2024 BCCRT 149Civil Resolution Tribunal of British Columbia, February 2024
  37. McDonald's to end AI drive-thru test with IBMCNBC, June 17, 2024
  38. NYC's AI Chatbot Tells Businesses to Break the LawThe Markup, March 29, 2024
  39. Mamdani to kill the NYC AI chatbot we caught telling businesses to break the lawThe Markup, January 30, 2026
  40. Klarna AI assistant handles two-thirds of customer service chats in its first monthKlarna, February 2024
  41. Klarna changes its AI tune and again recruits humans for customer serviceCX Dive, May 2025
  42. Commonwealth Bank backtracks on AI job cuts, apologises for 'error' as call volumes riseABC News, August 21, 2025
  43. AI-powered coding tool wiped out a software company's database in 'catastrophic failure'Fortune, July 23, 2025
  44. Cursor AI support bot hallucinated its own company policyThe Register, April 18, 2025
  45. Deloitte refunds Australian government over AI in reportThe Register, October 6, 2025
  46. Report: Starbucks scrapped an AI inventory tool and left a Seattle-area startup 'blindsided'GeekWire, 2026
  47. Klarna Group plc Form F-1 registration statementUS Securities and Exchange Commission, March 2025
  48. The state of AI: How organizations are rewiring to capture valueMcKinsey & Company, March 2025
  49. uReview: Scalable, Trustworthy GenAI for Code Review at UberUber Engineering, 2025
  50. Generative AI LensAWS Well-Architected, November 2025
  51. Microsoft Responsible AI Standard, v2: General RequirementsMicrosoft, June 2022
  52. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1)NIST, July 2024
  53. MLOps and GenAIOps for AI workloads on AzureMicrosoft Learn, Azure Well-Architected Framework
  54. A Radical Solution to Scale AI TechnologyHarvard Business Review, April 2020
  55. Shadow testsAmazon SageMaker AI documentation
  56. Levels of Autonomy for AI AgentsK. J. Kevin Feng, David W. McDonald and Amy X. Zhang, arXiv 2506.12469, June 2025
  57. Your AI Product Needs EvalsHamel Husain
  58. 2025: The State of Generative AI in the EnterpriseMenlo Ventures, December 2025
  59. How 100 Enterprise CIOs Are Building and Buying Gen AI in 2025Andreessen Horowitz, 2025
  60. Global Talent Shortage Reaches Turning Point as AI Skills Claim Top SpotYahoo Finance, 2026, on ManpowerGroup's Talent Shortage survey

Written by DigyAi Engineering from the systems we build and run. Every figure links to its public source, and every link and figure was checked on October 4, 2026. No client data appears in our insights.

Read next

All insights
  • Your own cases, stored in a golden-set archive, run through your system to a release gate whose board shows every segment against a threshold its owner signed in advance; one segment falls short and the gate holds, then the rerun clears the line and the release crosses a bridge to production.

    LLM and RAG engineering Guide

    LLM and AI Agent Evaluation: How to Prove a System Is Ready to Ship

    For CTOs, heads of AI and risk owners deciding whether an LLM application, RAG system or AI agent is ready to leave the pilot, and what evidence should back that decision.

    16 min read

  • An AI business case as a built place: a baseline from the systems of record and a set of graded cases feed one model, which raises three towers for the upside, base and downside cases. A glass deck marks the payback line, and the lit downside tower clearing it is the case the project is funded on.

    Economics and buying Playbook

    How to Calculate AI ROI Before You Commit Budget

    For CFOs, sponsors and CTOs deciding whether an AI system has earned its budget before the build begins.

    17 min read

  • Three AI agents send their calls through one lit router, which passes most of them to a fleet of small models and only a hard one to a large frontier model. Each agent has its own budget gauge; one has spent to its cap and a red barrier stops it, while the other two keep working.

    Economics and buying Playbook

    AI Inference Cost: How to Govern LLM and Agent Spend

    For CFOs, CTOs and FinOps leads deciding how to forecast, allocate and cap the recurring cost of LLM applications and AI agents in production.

    16 min read

Get in touch

Tell us what you are building.

Write it as big as you imagine it.