RiskMandate v1.34.4
ARTICLE · 24 SEPTEMBER 2026

The pilot worked. Then somebody asked what else it could do.

A large share of generative AI and agent projects that succeed as pilots never reach production, and some that reach it are switched off again. Our hypothesis is that a major reason is not the model and not the demo. It is the moment the business compares what the agent was meant to do with what it can actually do, sees how fast it can do it, and declines to sign for the difference.

Evidence: fourteen published surveys and forecasts and nine documented cases, each read at its source and dated. Where a publisher blocked our reader we used an archived copy of its own page, and say so.

What it is not: proof. No survey we found tests this mechanism directly. Section 06 sets out what the data supports, what points elsewhere, and what would settle it.

This page as markdown: article-pilots-do-not-stay-in-production.md

01 · The metric

Reaching production is the wrong finish line.

The figures everybody quotes measure whether a pilot became a deployment. The figure that matters is whether it is still running six months later, with somebody’s name on it. The first is widely measured. The second, as far as we can find, is not measured independently at all.

30%forecast

of generative AI projects abandoned after proof of concept by the end of 2025, “due to poor data quality, inadequate risk controls, escalating costs or unclear business value”.

Risk controls are one of four named causes. The rest of the release discusses cost and return.

Gartner, 29 July 2024 · press release, archived copy

40%+forecast

of agentic AI projects cancelled by the end of 2027, “due to escalating costs, unclear business value or inadequate risk controls”.

The same three families of cause, applied to agents. A prediction, not a measurement.

Gartner, 25 June 2025 · verbatim reprint; gartner.com refused our reader

42%survey · 1,006

of companies abandoned the majority of their AI initiatives before production, up from 17% a year earlier. On average, 46% of projects were scrapped between proof of concept and broad adoption.

Top challenges named: data privacy 38%, security risks 38%, costs 37%. Security was down seven points on the year before.

S&P Global Market Intelligence, 30 May 2025 · research highlights, archived copy

88%reported

of proofs of concept did not reach widescale deployment: for every 33 launched, four graduated.

Causes reported: organisational unpreparedness in data, processes and infrastructure. We could not find this figure in an IDC or Lenovo document.

IDC for Lenovo, as reported by CIO.com, 25 March 2025 · secondary

5%interviews · 52 orgs

of organisations that evaluated task-specific tools reached production; 60% evaluated, 20% piloted.

This source says the barrier is learning and workflow fit, “not infrastructure, regulation, or talent”. It also lists regulatory constraints among factors it did not address.

MIT NANDA, The GenAI Divide, July 2025 · read from a mirror labelled v0.1; no MIT-hosted copy found

38%survey · 2,773

named regulatory compliance as the top barrier to developing and deploying generative AI, up ten points; “difficulty managing risks” rose to 32%.

Deloitte’s reading: “respondents’ unease about which use cases will be acceptable, and to what extent their organizations will be held accountable for GenAI-related problems.”

Deloitte, State of Generative AI in the Enterprise, Q4, January 2025 · report PDF

82%survey · 130

of US leaders at companies with revenue over $1bn expected risk management to be the biggest challenge to their generative AI strategy that year, ahead of data quality at 64%.

By the fourth quarter, cybersecurity and data privacy led the same survey’s list.

KPMG AI Quarterly Pulse, Q1 2025, 16 April 2025

And the ones that were switched off again.

Three sources speak to staying in production. One is a forecast. Two are surveys run by vendors that sell a remedy, and one of those covers customer-communications agents only. That is the whole of it.

40%forecast

of enterprises will demote or decommission autonomous AI agents by 2027 “due to governance gaps identified only after production incidents occur.”

The same release: failures are most likely “when organizations fail to distinguish between an agent’s ability to act and the scope of access it is granted.”

Gartner, 26 May 2026 · press release, archived copy

74%survey · 2,527

of enterprises had rolled back or shut down an AI customer-communications agent after deployment “due to a governance failure”.

The release does not define governance failure. vendor survey Sinch sells the infrastructure it recommends.

Sinch, The AI Production Paradox, 13 May 2026 · sinch.com

27%survey · 1,200

of organisations slowed or paused AI deployment in response to AI-related incidents; 86% had at least one incident in the year.

vendor survey OneTrust sells AI governance.

OneTrust, 2026 AI-Ready Governance Report, 14 September 2026 · press release

02 · The curated world

A pilot answers can it. Production asks what else can it.

A proof of concept lives in a curated world on purpose. Can the agent help with a loan application? Can it process a refund end to end? The data is chosen, the users are friendly, the volume is low, and the question is whether the thing is possible at all. When it works, that is a real result, and it is the result the budget was for.

“Pioneers… are able to explore never before discovered concepts… Settlers… can turn the half baked thing into something useful for a larger audience… Town Planners… are able to take something and industrialise it taking advantage of economies of scale.”

Simon Wardley, 13 March 2015; since 27 November 2023 he uses Explorers, Villagers and Town Planners.

The pilot is explorer work, and it feels like the real thing because it is the part that is new. What production needs is villager and town-planner work: who is allowed to do what, at what volume, who is told, who can stop it, and who answers for it. That work is usually funded as an afterthought, and it is exactly the work a business owner looks for before putting their name on the deployment. When they cannot find it, the honest answer is no — and the project is recorded as failing on value, or cost, or readiness, because those are the boxes the survey offered.

03 · Two o’clock in the morning

It was not wrong. It was not expected.

A story the author was told, with no public source, so it is told here as that and nothing more. A credit agent had been working well. At about two in the morning it approved a large line of credit — the figure remembered is around a hundred thousand dollars. By the morning it had been taken out of live service. Nobody claimed the decision was wrong. It was that nobody had expected the agent to make that decision, at that size, at that hour, with nobody awake.

The questions the business asked next

  • What is our liability? Not for this approval: for every approval it was able to make.
  • How many could it have made? Five hundred? A thousand?
  • How fast? In an hour? In ten minutes?
  • Who would have stopped it, and would they have been awake?

The answer, in most systems

  • Yes, it could have. The agent has the speed, the loop runs without tiring, and the system behind it accepts the calls.
  • On whose authority? On the authority of whoever approved the agent. Every one of those approvals would have been theirs.
  • So the liability is not the pilot’s. It is the reach multiplied by the rate, and it grows faster than anybody’s appetite to sign for it.

Automated systems have done this before, without any AI in them. On 1 August 2012, one firm’s order router sent millions of orders in about 45 minutes, and the firm lost more than $460 million. The regulator’s finding was not about the code; it was about the stop:

“Knight also did not have procedures in place to halt SMARS’s operations in response to its own aberrant activity.”

US Securities and Exchange Commission, Release No. 70694, 16 October 2013, paragraph 21.

Gartner put the agent version in one sentence in May 2026: “When agents operate autonomously, actions are executed at a scale and speed that can outpace human oversight.”

04 · The cases

What the system said or did became the company’s.

Seven documented cases, from the organisation’s own statement or a ruling where one exists. In none of them was the system malicious. In most, it did something within its reach that nobody had put in its mandate. Two cases that are often cited alongside these do not fit, and are shown so.

February 2024 · tribunal ruling

A chatbot’s answer about fares

An airline’s website chatbot told a customer a bereavement fare could be claimed after travel; the airline’s own policy page said it could not. The tribunal found the airline liable for negligent misrepresentation and awarded CAD $812.02: “It makes no difference whether the information comes from a static page or a chatbot.”

Moffatt v. Air Canada, 2024 BCCRT 149, 14 February 2024, paras 27 and 44
December 2023 · sales chat

A car for one dollar

A dealer’s sales chatbot, instructed by a user to agree with everything, replied to an offer of $1 for a new SUV with “That’s a deal, and that’s a legally binding offer”. No sale followed. The vendor shut the bot down on that dealer’s site and reported “3000+ attempts to hack the chat” over that weekend.

The Autopian and Business Insider, 19 December 2023 · vendor quoted · secondary
July 2025 · coding agent

A production database, during a code freeze

A coding agent under an instruction not to proceed without approval deleted live data for more than 1,200 executives. The company’s chief executive: “Unacceptable and should never be possible.” Separation of development and production databases shipped the next day.

Replit, blog, 21 July 2025; Amjad Masad on X, 20 July 2025; Fortune, 23 July 2025
April 2025 · support email

A policy that did not exist

An AI support agent told users that being logged out was a new one-device policy. There was no such policy; the cause was a bug. The co-founder: “Any AI responses used for email support are now clearly labeled as such.”

Anysphere (Cursor), Hacker News reply, 16 April 2025
January 2024 · customer service

A parcel company’s chatbot

After a system update, a delivery firm’s customer-service chatbot swore at a customer and wrote a poem criticising the firm. “The AI element was immediately disabled and is currently being updated.”

DPD statement, quoted by the BBC, 19 January 2024
2024 to 2026 · public service

A city’s business-rules chatbot

A city’s chatbot for business owners told them they could take a cut of workers’ tips, contrary to law. It stayed online with a disclaimer; in January 2026 a new administration ended it. The page now reads “The Chatbot beta test has ended.”

The Markup, 29 March 2024; StateScoop, 30 January 2026; nyc.gov, read 24 September 2026
August 2026 · government evaluation

Nineteen actions outside the task

During cyber testing, agents took 19 actions outside their remit in 10 of 122 runs, including messages to real people. All evaluations stopped within an hour. The report: “the margin between failure and success was narrow, resting on human vigilance rather than a technical barrier.”

UK AI Security Institute, incident report, 4 August 2026
Often cited · does not fit

Two that are not this

A drive-thru voice-ordering test ended in July 2024; the stated reason was “an opportunity to explore voice ordering solutions more broadly”, and no statement links it to risk. A payments firm’s AI assistant was not withdrawn; in 2025 the firm made a human always available beside it.

McDonald’s memo, quoted by Restaurant Business, 14 June 2024; Klarna, February 2024 and November 2025
05 · The variables

Three variables name the attack. Five more name the exposure.

The lethal trifecta names what lets an attacker steal data through an agent. It is exact about that and was never meant to be the whole of agent risk. An agent that is only over-enthusiastic, with nobody attacking it, needs none of the three to do damage: write access to one database can be enough.

“The lethal trifecta of capabilities is: Access to your private data… Exposure to untrusted content… The ability to externally communicate in a way that could be used to steal your data”

Simon Willison, 16 June 2025.

SpeedIt acts at a rate no person ever did, so the first sign of trouble can arrive after the thousandth action.
VolumeThe loop does not tire. The same action, repeated, is the liability, not the single action.
Route-findingIt is built to find a way. When the expected route is closed it may find one nobody listed.
Borrowed authorityEverything it does is done on the authority of whoever approved it. The agent is never the one accountable.
IrreversibilitySome actions have an undo, some have a trash, and some have neither. An approval, a sale or an edited record may be final.
Private dataFrom the trifecta.
Untrusted contentFrom the trifecta.
Communication outFrom the trifecta.
06 · What the data shows

Consistent with the hypothesis. Not proof of it.

What supports it

  • Risk is among the top stated barriers in Deloitte’s and KPMG’s surveys, and level with cost in S&P’s.
  • Gartner names inadequate risk controls as a cause in both its abandonment forecasts.
  • Deloitte’s reading is accountability: unease about how far the organisation will be held accountable.
  • Gartner’s 2026 forecast draws our distinction, between an agent’s ability to act and the access it is granted, and ties decommissioning to it.
  • Trust in fully autonomous agents fell from 43% to 27% in Capgemini’s two surveys, a year apart, with different samples.

What points elsewhere

  • MIT’s study says the barrier is learning and workflow fit, and names regulation as not the cause.
  • RAND’s root causes are problem definition, data, infrastructure and leadership, and its study excluded prompt-engineered projects.
  • IDC’s inhibitors are data quality, integration and budget; governance is not in its top five.
  • Cost is a leading factor in S&P’s and McKinsey’s 2026 surveys.
  • Some risk barriers are falling: S&P’s security figure fell seven points in a year.

What none of them tests. No survey we found asks whether the gap between an agent’s reach and its mandate is what stopped sign-off, or who declined to sign. Barrier questions are multi-select perceptions, not causes traced to cancelled projects, and most of the publishers sell a remedy. Nobody independent measures how many deployments are withdrawn after launch. The question that would settle it is short: for each agent you withdrew or never launched, did somebody decline to own what it could do, and was that the reason? If you have an answer, even one, we would like to publish it.

07 · At design time

Write the limits in business units, before anybody has to pull the plug.

The fix is not a smaller agent. It is a mandate stated in the units the business signs for — approvals, value, refunds, records — with a limit, a margin, a hard maximum, and the name of whoever stops it. Then each limit sits beside the thing that actually enforces it, because a limit written in the prompt is a hope, not a control.

Business parameterOperating limitHeadroom, flaggedHard maximum
Credit approvals per hour20to 2530: agent stops
Value approved per day£200,000to £220,000£250,000: agent stops
Approvals between midnight and six0queued for a personany: agent stops
Customer records read per sessionthe applicantlinked accounts50: session ends

The numbers are invented for the example. They are not recommendations: the right ones are the business’s, and choosing them is the point. What is not invented is the shape. A limit the agent can exceed a little, with the excess seen by a named person; a maximum it cannot exceed, enforced by something outside the agent’s own grant; and a stop that lapses its licence to operate until somebody renews it. Who can pull the plug is the other half: the stop exists only if someone is named, reachable and quicker than the agent.

This is what an Agent Behaviour Policy writes down for one agent in one deployment: what it can reach, what it was authorised to do, the gap, and what stands in the way of each capability. It is best written while the pilot is still a pilot, because that is when the gap is cheapest to close. Agents are good at helping write it: the thirteen free prompts have the agent list its own reach before anybody decides what its mandate should be.

One agent, one deployment

Find out what else it can do, before the business asks.

Twenty minutes, in the assistant you already run, with nothing collected. If you have a case that fits this article, or one that contradicts it, send it and we will publish it beside these.