Agents Will Be Everywhere. Most Agent Projects Will Be Cancelled.

Agentic AI projects following contrasting paths towards successful deployment or cancellation due to rising costs, unclear value and late risk controls.

Quick answer: Gartner forecasts that 40% of enterprise applications will carry task-specific agents by the end of this year, and that over 40% of agentic AI projects will be cancelled by the end of next. Both are probably right, and the space between them is the most useful thing said about agents this year.

Two Gartner forecasts, placed side by side, look like a contradiction.

The first: 40% of enterprise applications will include task-specific AI agents by the end of 2026, up from fewer than 5% in 2025. The second: more than 40% of agentic AI projects will be cancelled by the end of 2027, attributed to escalating costs, unclear business value and inadequate risk controls.

They are not in conflict. They describe different populations. The first is about features shipped inside products that vendors build once and sell many times. The second is about bespoke projects that individual organisations attempt themselves. Rapid diffusion of a capability alongside a high failure rate among in-house implementations is the ordinary signature of a technology in its scaling phase.

Is adoption real, and uneven?

McKinsey’s early-September research describes a two-speed race. Among organisations with more than $1 billion in annual revenue, 40% report they are now scaling AI agents, up from 27% a year earlier. Across the broader sample, 23% are scaling an agentic system somewhere in the enterprise while a further 39% are still experimenting.

So the demand is genuine and the movement from pilot to production is happening. The open question is what proportion survives contact with a profit-and-loss account.

What are the three stated causes, and do they deserve to be taken seriously?

Escalating costs. Agent loops consume tokens non-linearly. A workflow that costs pennies as a single call can cost pounds when an agent plans, calls tools, evaluates and retries. OpenAI reported its own researchers running 3.1 agent-workdays per human workday, with median spend above $600 per researcher per day at API prices. Most business cases are built on per-call arithmetic and then meet per-task reality.

Unclear business value. The familiar problem in a new costume: capability acquired before a measurable baseline was identified. If you cannot state what the process costs today, you cannot demonstrate that the agent improved it.

Inadequate risk controls. The most expensive of the three, because it tends to surface late. A system that works is stopped at the security, legal or compliance gate after the money has been spent, not before.

Is agent washing part of the explanation?

Gartner also points to a supply-side problem. Vendors have rebranded existing products as agentic without substantive agentic capability, and Gartner’s estimate is that only around 130 of the thousands of vendors claiming the label are the real thing.

That figure, if roughly right, explains a meaningful share of the projected cancellations on its own. A buyer selecting from a market where most labels are inaccurate will sometimes buy something that cannot do what the category name implies. The project is then cancelled — correctly — but the failure was in procurement, not in the technology.

A cancellation rate is not a verdict on a technology. It is a measurement of the gap between what was promised, what was bought, and what was actually needed.

What’s the honest counter-argument?

It is worth resisting the temptation to read the cancellation figure as proof that agents do not work.

A 40% cancellation rate would be unremarkable for any category of ambitious internal IT project, and Gartner’s own framing is that most current agentic efforts are early-stage experiments and proofs of concept. Cancelling an experiment that has answered its question is a success, not a failure. The statistic is only alarming to organisations that were treating experiments as commitments — which is a governance problem rather than a technical one.

There is also a reasonable argument that the forecast is self-defeating in the useful sense: published widely enough, it changes the behaviour it predicts, because boards start asking harder questions before approving the next pilot.

What separates the projects that survive?

A bounded scope with a measured baseline. The strongest returns continue to come from high-volume, repetitive work with an unambiguous success criterion and a known current cost. That is unglamorous and it is where the evidence points.

Cost modelled per completed task. Not per call, and not per token. Model the loop, including retries and failures, and set a ceiling that stops a runaway before it reaches the invoice.

Risk controls designed in, not appended. Bring security, legal and compliance into the design conversation. The gate they represent is cheaper to pass in week two than in month nine.

Human checkpoints where being wrong is expensive. Full autonomy is a choice to be made per decision type, on the basis of consequence, rather than adopted wholesale because the technology permits it.

What’s the honest summary?

The two forecasts together describe a market where the capability is arriving regardless of what any individual organisation decides, and where a large share of deliberate attempts to build with it will be abandoned.

The differentiator is not ambition or access to the technology. Both are widely available. It is the discipline of scoping to a measurable problem, costing the loop rather than the call, and involving the people who can stop the project before the money is spent rather than after.

None of that is novel advice. It is the same advice that applied to every previous wave of enterprise software, which is precisely why it keeps being ignored.

Note on currency: figures are quoted in US dollars as published in the underlying research.

Key takeaways

  • Gartner’s two forecasts aren’t contradictory: 40% of enterprise apps will ship with built-in agents by end of 2026 (vendor-built features), while over 40% of bespoke, in-house agentic projects will be cancelled by end of 2027. Different populations, both plausible.
  • McKinsey found a two-speed adoption race: 40% of $1bn+ revenue organisations are now scaling AI agents (up from 27% a year earlier), versus 23% scaling and 39% still experimenting across the broader sample.
  • The three stated causes of cancellation are escalating costs (agent loops consume tokens non-linearly, OpenAI reported median spend above $600/researcher/day), unclear business value (no measurable baseline to compare against), and risk controls surfacing late, after the money’s already spent.
  • “Agent washing” compounds the problem: Gartner estimates only around 130 of thousands of vendors claiming the “agentic” label actually have the substantive capability, so some cancellations are a procurement failure, not a technology failure.
  • A 40% cancellation rate is unremarkable for ambitious internal IT projects generally, and cancelling an experiment that already answered its question is a success, not a failure, the real problem is treating experiments as commitments.
  • What separates surviving projects: a bounded scope with a measured baseline, cost modelled per completed task rather than per call or token, risk controls designed in from the start rather than appended late, and human checkpoints placed where being wrong is genuinely expensive.

FAQs

Do Gartner’s two agentic AI forecasts contradict each other?
No. The 40% enterprise-app figure describes agent features built into vendor products and sold at scale. The 40% cancellation figure describes bespoke, in-house agentic projects individual organisations build themselves. Rapid feature diffusion alongside a high failure rate in custom builds is normal for a technology in its scaling phase.

Why are agentic AI projects getting cancelled?
Gartner points to three causes: escalating costs (agent loops consume tokens non-linearly, so per-call cost estimates badly understate per-task reality), unclear business value (no measurable baseline existed before the project started), and inadequate risk controls (the project gets stopped at a security, legal, or compliance gate after the budget is already spent).

What is “agent washing”?
It’s vendors rebranding existing, non-agentic products with the “agentic” label without substantive agentic capability. Gartner estimates only around 130 of the thousands of vendors making this claim actually have the real capability, which means some project cancellations are really procurement failures rather than a sign the technology doesn’t work.

Is a 40% cancellation rate actually bad?
Not necessarily. It’s an unremarkable rate for ambitious internal IT projects generally, and Gartner’s own framing is that most current agentic efforts are early-stage experiments. Cancelling an experiment that has already answered its question is a successful outcome, not a failed one.

What do successful agentic AI projects have in common?
Four things: a bounded scope with a known, measured baseline cost; cost modelled per completed task rather than per call or token; risk controls (security, legal, compliance) designed in from the start rather than added later; and human checkpoints placed specifically where a wrong decision would be expensive.

Have a project in mind? Let’s get to work.

Let’s chat about how we can help you. Fill in the details and we’ll get back to you as soon we can.