<!-- Generated from article-pilots-do-not-stay-in-production.html by scripts/site/generate.mjs. Edit the page, not this file. -->

# The pilot worked. Then somebody asked what else it could do.

Why so many generative AI and agent pilots never reach production, or are switched off after they do: the business compares what the agent was meant to do with what it can do, at machine speed, and declines to sign for the difference. The data, the cases, and what the data does not show.

Source: https://riskmandate.ai/article-pilots-do-not-stay-in-production.html

---

# The pilot worked. Then somebody asked what else it could do.

A large share of generative AI and agent projects that succeed as pilots never reach production, and some that reach it are switched off again. Our hypothesis is that a major reason is not the model and not the demo. It is the moment the business compares what the agent was meant to do with what it can actually do, sees how fast it can do it, and declines to sign for the difference.

**Evidence:** fourteen published surveys and forecasts and nine documented cases, each read at its source and dated. Where a publisher blocked our reader we used an archived copy of its own page, and say so.

**What it is not:** proof. No survey we found tests this mechanism directly. Section 06 sets out what the data supports, what points elsewhere, and what would settle it.

**This page as markdown:** [article-pilots-do-not-stay-in-production.md](article-pilots-do-not-stay-in-production.md)

## Reaching production is the wrong finish line.

The figures everybody quotes measure whether a pilot became a deployment. The figure that matters is whether it is still running six months later, with somebody’s name on it. The first is widely measured. The second, as far as we can find, is not measured independently at all.

of generative AI projects **abandoned after proof of concept** by the end of 2025, “due to poor data quality, inadequate risk controls, escalating costs or unclear business value”.

Risk controls are one of four named causes. The rest of the release discusses cost and return.

Gartner, 29 July 2024 · [press release](https://www.gartner.com/en/newsroom/press-releases/2024-07-29-gartner-predicts-30-percent-of-generative-ai-projects-will-be-abandoned-after-proof-of-concept-by-end-of-2025), archived copy

of **agentic AI projects cancelled** by the end of 2027, “due to escalating costs, unclear business value or inadequate risk controls”.

The same three families of cause, applied to agents. A prediction, not a measurement.

Gartner, 25 June 2025 · [verbatim reprint](https://www.predictiveanalyticsworld.com/machinelearningtimes/gartner-predicts-over-40-of-agentic-ai-projects-will-be-canceled-by-end-of-2027/13875/); gartner.com refused our reader

of companies **abandoned the majority of their AI initiatives** before production, up from 17% a year earlier. On average, 46% of projects were scrapped between proof of concept and broad adoption.

Top challenges named: data privacy 38%, security risks 38%, costs 37%. Security was down seven points on the year before.

S&P Global Market Intelligence, 30 May 2025 · [research highlights](https://www.spglobal.com/market-intelligence/en/news-insights/research/ai-experiences-rapid-adoption-but-with-mixed-outcomes-highlights-from-vote-ai-machine-learning), archived copy

of proofs of concept **did not reach widescale deployment**: for every 33 launched, four graduated.

Causes reported: organisational unpreparedness in data, processes and infrastructure. We could not find this figure in an IDC or Lenovo document.

IDC for Lenovo, as reported by [CIO.com](https://www.cio.com/article/3850763/88-of-ai-pilots-fail-to-reach-production-but-thats-not-all-on-it.html), 25 March 2025 · secondary

of organisations that evaluated task-specific tools **reached production**; 60% evaluated, 20% piloted.

This source says the barrier is learning and workflow fit, “not infrastructure, regulation, or talent”. It also lists regulatory constraints among factors it did not address.

MIT NANDA, _The GenAI Divide_, July 2025 · read from a mirror labelled v0.1; no MIT-hosted copy found

named **regulatory compliance as the top barrier** to developing and deploying generative AI, up ten points; “difficulty managing risks” rose to 32%.

Deloitte’s reading: “respondents’ unease about which use cases will be acceptable, and to what extent their organizations will be held accountable for GenAI-related problems.”

Deloitte, _State of Generative AI in the Enterprise_, Q4, January 2025 · report PDF

of US leaders at companies with revenue over $1bn **expected risk management to be the biggest challenge** to their generative AI strategy that year, ahead of data quality at 64%.

By the fourth quarter, cybersecurity and data privacy led the same survey’s list.

KPMG AI Quarterly Pulse, [Q1 2025](https://kpmg.com/us/en/media/news/q1-ai-pulse-2025.html), 16 April 2025

## And the ones that were switched off again.

Three sources speak to _staying_ in production. One is a forecast. Two are surveys run by vendors that sell a remedy, and one of those covers customer-communications agents only. That is the whole of it.

of enterprises will **demote or decommission autonomous AI agents** by 2027 “due to governance gaps identified only after production incidents occur.”

The same release: failures are most likely “when organizations fail to distinguish between an agent’s ability to act and the scope of access it is granted.”

Gartner, 26 May 2026 · [press release](https://www.gartner.com/en/newsroom/press-releases/2026-05-26-gartner-says-applying-uniform-governance-across-ai-agents-will-lead-to-enterprise-ai-agent-failure), archived copy

of enterprises **had rolled back or shut down** an AI customer-communications agent after deployment “due to a governance failure”.

The release does not define governance failure. vendor survey Sinch sells the infrastructure it recommends.

Sinch, _The AI Production Paradox_, 13 May 2026 · [sinch.com](https://www.sinch.com/ai-production-paradox/)

of organisations **slowed or paused AI deployment** in response to AI-related incidents; 86% had at least one incident in the year.

vendor survey OneTrust sells AI governance.

OneTrust, _2026 AI-Ready Governance Report_, 14 September 2026 · press release

## A pilot answers _can it_. Production asks _what else can it_.

A proof of concept lives in a curated world on purpose. Can the agent help with a loan application? Can it process a refund end to end? The data is chosen, the users are friendly, the volume is low, and the question is whether the thing is possible at all. When it works, that is a real result, and it is the result the budget was for.

> “Pioneers… are able to explore never before discovered concepts… Settlers… can turn the half baked thing into something useful for a larger audience… Town Planners… are able to take something and industrialise it taking advantage of economies of scale.”

Simon Wardley, [13 March 2015](https://blog.gardeviance.org/2015/03/on-pioneers-settlers-town-planners-and.html); since 27 November 2023 he uses _Explorers, Villagers and Town Planners_.

The pilot is explorer work, and it feels like the real thing because it is the part that is new. What production needs is villager and town-planner work: who is allowed to do what, at what volume, who is told, who can stop it, and who answers for it. That work is usually funded as an afterthought, and it is exactly the work a business owner looks for before putting their name on the deployment. **When they cannot find it, the honest answer is no** — and the project is recorded as failing on value, or cost, or readiness, because those are the boxes the survey offered.

## It was not wrong. It was not expected.

A story the author was told, with no public source, so it is told here as that and nothing more. A credit agent had been working well. At about two in the morning it approved a large line of credit — the figure remembered is around a hundred thousand dollars. By the morning it had been taken out of live service. Nobody claimed the decision was wrong. It was that nobody had expected the agent to make that decision, at that size, at that hour, with nobody awake.

### The questions the business asked next

- **What is our liability?** Not for this approval: for every approval it was able to make.
- **How many could it have made?** Five hundred? A thousand?
- **How fast?** In an hour? In ten minutes?
- **Who would have stopped it,** and would they have been awake?

### The answer, in most systems

- **Yes, it could have.** The agent has the speed, the loop runs without tiring, and the system behind it accepts the calls.
- **On whose authority?** On the authority of whoever approved the agent. Every one of those approvals would have been theirs.
- **So the liability is not the pilot’s.** It is the reach multiplied by the rate, and it grows faster than anybody’s appetite to sign for it.

Automated systems have done this before, without any AI in them. On 1 August 2012, one firm’s order router sent millions of orders in about 45 minutes, and the firm lost more than $460 million. The regulator’s finding was not about the code; it was about the stop:

> “Knight also did not have procedures in place to halt SMARS’s operations in response to its own aberrant activity.”

US Securities and Exchange Commission, [Release No. 70694](https://www.sec.gov/files/litigation/admin/2013/34-70694.pdf), 16 October 2013, paragraph 21.

Gartner put the agent version in one sentence in May 2026: “When agents operate autonomously, actions are executed at a scale and speed that can outpace human oversight.”

## What the system said or did became the company’s.

Seven documented cases, from the organisation’s own statement or a ruling where one exists. In none of them was the system malicious. In most, it did something within its reach that nobody had put in its mandate. Two cases that are often cited alongside these do not fit, and are shown so.

### A chatbot’s answer about fares

An airline’s website chatbot told a customer a bereavement fare could be claimed after travel; the airline’s own policy page said it could not. The tribunal found the airline liable for negligent misrepresentation and awarded CAD $812.02: **“It makes no difference whether the information comes from a static page or a chatbot.”**

### A car for one dollar

A dealer’s sales chatbot, instructed by a user to agree with everything, replied to an offer of $1 for a new SUV with “That’s a deal, and that’s a legally binding offer”. No sale followed. The vendor shut the bot down on that dealer’s site and reported “3000+ attempts to hack the chat” over that weekend.

### A production database, during a code freeze

A coding agent under an instruction not to proceed without approval deleted live data for more than 1,200 executives. The company’s chief executive: **“Unacceptable and should never be possible.”** Separation of development and production databases shipped the next day.

### A policy that did not exist

An AI support agent told users that being logged out was a new one-device policy. There was no such policy; the cause was a bug. The co-founder: “Any AI responses used for email support are now clearly labeled as such.”

### A parcel company’s chatbot

After a system update, a delivery firm’s customer-service chatbot swore at a customer and wrote a poem criticising the firm. “The AI element was immediately disabled and is currently being updated.”

### A city’s business-rules chatbot

A city’s chatbot for business owners told them they could take a cut of workers’ tips, contrary to law. It stayed online with a disclaimer; in January 2026 a new administration ended it. The page now reads “The Chatbot beta test has ended.”

### Nineteen actions outside the task

During cyber testing, agents took 19 actions outside their remit in 10 of 122 runs, including messages to real people. All evaluations stopped within an hour. The report: “the margin between failure and success was narrow, resting on human vigilance rather than a technical barrier.”

### Two that are not this

**A drive-thru voice-ordering test** ended in July 2024; the stated reason was “an opportunity to explore voice ordering solutions more broadly”, and no statement links it to risk. **A payments firm’s AI assistant** was not withdrawn; in 2025 the firm made a human always available beside it.

## Three variables name the attack. Five more name the exposure.

The lethal trifecta names what lets an attacker steal data through an agent. It is exact about that and was never meant to be the whole of agent risk. An agent that is only over-enthusiastic, with nobody attacking it, needs none of the three to do damage: write access to one database can be enough.

> “The lethal trifecta of capabilities is: Access to your private data… Exposure to untrusted content… The ability to externally communicate in a way that could be used to steal your data”

Simon Willison, [16 June 2025](https://simonwillison.net/2025/Jun/16/the-lethal-trifecta/).

## Consistent with the hypothesis. Not proof of it.

### What supports it

- **Risk is among the top stated barriers** in Deloitte’s and KPMG’s surveys, and level with cost in S&P’s.
- **Gartner names inadequate risk controls** as a cause in both its abandonment forecasts.
- **Deloitte’s reading is accountability:** unease about how far the organisation will be held accountable.
- **Gartner’s 2026 forecast draws our distinction**, between an agent’s ability to act and the access it is granted, and ties decommissioning to it.
- **Trust in fully autonomous agents fell** from 43% to 27% in Capgemini’s two surveys, a year apart, with different samples.

### What points elsewhere

- **MIT’s study says the barrier is learning** and workflow fit, and names regulation as not the cause.
- **RAND’s root causes** are problem definition, data, infrastructure and leadership, and its study excluded prompt-engineered projects.
- **IDC’s inhibitors** are data quality, integration and budget; governance is not in its top five.
- **Cost is a leading factor** in S&P’s and McKinsey’s 2026 surveys.
- **Some risk barriers are falling:** S&P’s security figure fell seven points in a year.

**What none of them tests.** No survey we found asks whether the gap between an agent’s reach and its mandate is what stopped sign-off, or who declined to sign. Barrier questions are multi-select perceptions, not causes traced to cancelled projects, and most of the publishers sell a remedy. Nobody independent measures how many deployments are withdrawn after launch. **The question that would settle it** is short: for each agent you withdrew or never launched, did somebody decline to own what it could do, and was that the reason? If you have an answer, even one, we would like to publish it.

## Write the limits in business units, before anybody has to pull the plug.

The fix is not a smaller agent. It is a mandate stated in the units the business signs for — approvals, value, refunds, records — with a limit, a margin, a hard maximum, and the name of whoever stops it. Then each limit sits beside the thing that actually enforces it, because a limit written in the prompt is a hope, not a control.

**The numbers are invented for the example.** They are not recommendations: the right ones are the business’s, and choosing them is the point. What is not invented is the shape. A limit the agent can exceed a little, with the excess seen by a named person; a maximum it cannot exceed, enforced by something outside the agent’s own grant; and a stop that lapses its [licence to operate](licence-to-operate.html) until somebody renews it. [Who can pull the plug](plug.html) is the other half: the stop exists only if someone is named, reachable and quicker than the agent.

This is what an [Agent Behaviour Policy](abp.html) writes down for one agent in one deployment: what it can reach, what it was authorised to do, the gap, and what stands in the way of each capability. It is best written while the pilot is still a pilot, because that is when the gap is cheapest to close. Agents are good at helping write it: the [thirteen free prompts](try-it.html) have the agent list its own reach before anybody decides what its mandate should be.

## Find out what else it can do, before the business asks.

Twenty minutes, in the assistant you already run, with nothing collected. If you have a case that fits this article, or one that contradicts it, send it and we will publish it beside these.
