Hand a team a high-capacity battery and no circuit diagram, and the rational thing to do is wire it to everything within reach. You get five bulbs at half strength. A weak, flickering glow where there could have been one clear light.
That is the state of a great deal of public sector AI investment. The capability is genuinely there. The charge that could have powered one strategic outcome bleeds out across five processes nobody ever ranked against each other.
The evidence for this is no longer thin. In June 2025 Gartner predicted that more than four in ten agentic AI projects would be cancelled by the end of 2027, and the reasons it gave had almost nothing to do with the technology: escalating costs, inadequate risk controls, no clear business value. In August 2025 MIT's NANDA study of enterprise AI found that roughly ninety-five per cent of pilots delivered no measurable impact, and located the cause not in the models but in the space between the capability and the organisation around it. A year on, both read less like forecasts and more like descriptions of the present.
So the usual explanations need retiring. This is not a skills gap. It is not procurement being slow. It is a failure at the join, in the space between what an institution intends and what it actually buys, and it has three distinct mechanisms. Each one is fixable. None of them is fixed by a better model.
Mechanism one: nobody checked the axioms
The most expensive mistakes in digital delivery do not come from bad execution. They come from unquestioned assumptions baked in at the start.
The pattern is always recognisable. A policy gets digitised exactly as it existed on paper, inefficiencies and all. A dashboard measures what was always measured rather than what matters. A business case is approved because the framing feels familiar, not because the underlying need was ever tested.
Most of us in digital and data spend our days reasoning by analogy. We copy what worked last quarter. We adopt frameworks because they were validated somewhere else. Often that is sensible: it reduces uncertainty and it is efficient. Until it is not.
Senior roundtables carry a quiet risk of becoming echo chambers of best-practice tourism. Someone cites a success from another department and it is quickly treated as a portable blueprint. The assumption buried in that move, that two environments sharing a label also share the same underlying conditions, is where value starts to leak. In government especially, the infrastructure beneath the work changes what is possible: policy, legislation, operational reality, institutional memory, accountability. Contexts are not interchangeable.
Three prompts catch this before commitment.
Evidence the need before requirements are written. The trap is treating the request as the need. Requests are solutions wearing a disguise. Before any requirement is drafted, someone should be able to finish one sentence with confidence: the user is trying to do something, cannot, because of something, resulting in something. If that sentence will not close, there is no stable foundation yet. Name the constraint set early too: statute, operational reality, risk thresholds, legacy dependencies, funding and governance shape. That is not pessimism. It is the difference between a viable path and an optimistic one. And ask the counterfactual. If the do-nothing scenario is acceptable, the work may not be justified. If it is unacceptable, you now have the anchor for prioritisation.
Work in probabilities, not certainties. Benefits are forecasts, conditional on adoption, behaviour change, process change and data quality. Stating them as certainties does not make them more likely; it makes them harder to manage when they do not arrive. Replace point estimates with ranges. Not "this will save five million" but "our estimate is two to six million, and here are the conditions required to land in the upper half." Separate benefit potential from benefit probability. A proposal can have high upside and low likelihood at the same time.
Seek disproof before approval. Business cases are sometimes built to win approval rather than to survive reality. A thirty-minute pre-mortem can save a year of sunk cost. It is twelve months from now and this failed: what happened? You are not predicting failure, you are surfacing fragility while it is still cheap. Which assumption, if wrong, kills the value? What evidence would a sceptic demand?
None of this is new. What is new is the speed at which a flawed assumption now propagates. If the data is biased, the objective is wrong or the user need is misunderstood, AI will not pause to query it. It will operationalise it. In a pre-AI environment a misaligned investment failed slowly and you had time to notice. In an AI-enabled environment you do not.
Mechanism two: the words stopped pointing at anything
In 1981 Jean Baudrillard described a map so detailed it replaced the territory entirely. The cartographers kept updating the map while the territory beneath it quietly collapsed. He called the result the simulacrum: not a representation of reality but a copy that has lost its original, a sign with no referent left.
He was writing about postmodern culture. He might as well have been writing about enterprise AI procurement.
"Agent." "Copilot." "AI-powered." "Intelligent." "Frontier." Each arrived with real meaning. Each has been hollowed out through overuse, misapplication, and the entirely rational commercial incentive to attach premium vocabulary to commodity products. Gartner has a name for the practice: agentwashing, relabelling existing automation, rule-based workflows or simple chatbot functionality as agentic AI to capture a premium price and a share of executive attention. Gartner also places AI agents at the Peak of Inflated Expectations on its Hype Cycle, which makes the drop into disillusionment a scheduled event rather than a risk. The organisations that built strategies on the map rather than the territory will feel it hardest.
This matters most when the vendor is large, because scale makes branding structurally significant. When a dominant productivity suite is rebranded around a premium term while the underlying product is materially unchanged, that redefinition propagates through strategy decks and procurement frameworks across the public and private sectors at once. Nobody has to be dishonest for this to happen. The words simply arrive faster than the scrutiny.
And it is a value problem, not a linguistic one. McKinsey's research puts the share of organisations seeing real financial return from AI at scale at around 5.5 per cent. IBM's Institute for Business Value finds enterprise AI return depends overwhelmingly on whether implementation was anchored in genuine business need from the outset. Deloitte puts payback at two to four years, and only for programmes where strategic intent and provider capability were aligned before any money moved. The pattern is consistent. The simulacrum costs money, and significantly.
Here is the connection I want to draw, because it sits across both mechanisms and I have not seen it made directly. A hollow word is an unexamined axiom with a price tag attached. "We need an agent" is not a requirement. It is an assumption wearing procurement language. It fails at exactly the same moment as every other unexamined axiom: before commitment, invisibly, in a room where everyone nodded.
Which means the counter is the same in both cases. When a vendor says agent, the question is: agent in what sense? Autonomous under what conditions? With what failure modes? Integrated with which systems? Trained on what data? At what cost? These are not hostile questions. They are the questions that separate value realisation from value leakage, and they have to be asked upstream of procurement, before commitments are made and contracts signed.
Baudrillard's cartographers did not set out to replace the territory. They were simply very good at drawing maps. Commercial incentives do the rest. The map is very well designed. Point at the territory anyway.
Mechanism three: a circuit diagram with no current
It helps to be precise about what an AI system actually is, because the vocabulary problem above hides the architecture underneath.
A working agent is a stack. A generative model sits at the base. Above it, retrieval and contextual memory decide what it can know. Higher up, specialised tools and delegated subagents decide what it can do. At the top, orchestration and observability decide whether it acted safely. Production value depends on the quality of the whole stack, not on the cleverness of the model at the bottom of it.
Read that stack against a procurement pattern and the leak becomes obvious. Large institutions buy the bottom layer, which is the layer with a price list, and underinvest in every layer above it, which is the part that determines whether anything works. Most of the discipline agentic AI demands lives in those upper layers. That is exactly where governance and demand shaping operate, and exactly where most organisations are thinnest.
Enterprise architecture was meant to prevent this. It supplies the approved integration patterns, the data governance boundaries, the technical conditions under which an investment makes sense. It is the circuit diagram, and it is genuinely valuable. But it carries a structural blind spot: it maps the wiring and it does not see the room. The business architecture layer sits at the top of the conventional stack, and in many large institutions it stays theoretical, populated by governance documents rather than practitioners, describing the business rather than engaging with it. In an environment where capability spreads faster than any governance cycle can track it, a theoretical top layer is a liability.
A diagram does not direct current. Scattered local enthusiasm only becomes legible as demand when someone reads AI requests against strategic objectives and clusters them by the pattern of work they represent. That is demand intelligence, and it is worth being blunt about what it is not. It is not a system you can buy. It is a capability you cultivate and a relationship someone has to hold. Doing the same thing with standard procurement cycles will simply yield the same scattered, failed pilots.
There is a related instinct worth intercepting. Faced with a bounded problem, organisations reach for the largest frontier model available, on the assumption that the smartest system is always required. A frontier model is a vast generalist engine. Pointing it at routine document processing is both a risk and a waste of compute. The technical answer is usually distillation: extract the specific reasoning capability and bake it into a smaller, locally governed system. The institutional answer is someone whose job it is to stop the organisation renting intelligence blindly, so the state builds specific, sovereign tools instead.
Why the third mechanism is the one that compounds
Satya Nadella has made the argument that models will commoditise and the scarce asset becomes the learning loop a firm builds on top of them, where human capital and what he calls token capital compound together. Human capital is the judgement, relationships and pattern recognition your people carry. Token capital is the AI capability the institution builds and owns. His claim is that the first grows in value as the machines improve, because human agency sets the goals, connects the dots across domains and decides which patterns matter. Strip the direction out and you have compute running in circles.
Gennaro Cuofano mapped the same thread to what he calls harness theory: the model is the commodity, and the harness wrapped around it is the actual moat.
Both are right about the architecture. Both leave the same thing out. Compounding does not start on its own, and a harness requires an architect. A learning loop needs someone to close it: to notice that a capability built in one directorate answers a need forming in another, to hold the connection open long enough for it to be used, to decide which patterns are worth compounding at all. The technology can do the capturing. Agents can already absorb years of meetings and decisions. What they cannot do is decide what is worth keeping. An agent captures what it is pointed at, which makes the act of pointing the actual work.
Nadella describes the physics. He does not name the orchestrator. In most organisations, nobody has.
Read from inside government, this lands with a different weight. The private sector version of the worry is competitive: cede your institutional knowledge to a handful of models and a rival with a better learning loop takes the ground from under you. The public sector version is closer to constitutional. It concerns the capacity of the state to know what it knows, to act on it accountably, and to explain itself afterwards. When a department's accumulated judgement leaks into a general model that nobody in the building controls, what goes is not competitive advantage. It is a piece of how the state governs itself. On that reading, the orchestration layer stops being a convenience and becomes a condition of accountable government.
What actually gets in the way
There is an uncomfortable irony here, and it explains why so many organisations have this capability on paper and not in practice.
My Henley professor of informatics, Kecheng Liu, named the shape of it precisely: the inverted onion. Low-value work crowds out strategic work. The person who should be shaping demand becomes the intake mechanism. An AI request arrives, nobody is sure who should assess it or what questions to ask, and it lands in an inbox. Three weeks of triage later, the strategic conversation has not happened. As AI use case requests multiply, the inversion accelerates.
Think of what happened to hospital emergency departments before clinical triage existed. Every patient arrived at the same door and joined the same queue, so the most skilled clinicians spent their hours on the most straightforward presentations simply because those arrived first. Triage did not replace doctors. It redirected them, and the system breathed.
So the honest sequence is this. Design the intake before you scale the ambition.
Start by mapping the process as it actually operates rather than as it was designed to operate. Most people discover, doing this honestly, that intake has never been formally mapped at all; it exists as institutional habit maintained by individual effort. You cannot automate a process you have not yet seen clearly. Then map the value: which functions generate the highest volume of requests, and which requests would produce the most downstream value if handled strategically. Audit what you already own before buying anything, because the workflow logic sitting idle in your current service management platform will usually carry the first routing rule without a business case. Build the playbook, defining precisely what qualifies for automatic routing, what requires human engagement, and what markers indicate elevated risk: data classification, third-party involvement, cross-boundary implications, executive sensitivity, live procurement. Every high-risk request that clears intake should feed the AI risk register. A playbook that maps triggers to intervention points gives any automated layer its decision logic; without it, automation is faster randomness. Finally, heat-map where requests actually stall, whether at governance review, security assurance or procurement sign-off, and be rigorous about the distinction between friction that routing can mitigate and friction that requires process redesign upstream. Conflating those two produces a faster version of the same blockage.
That sequence assumes something many public sector organisations cannot currently offer, which is the capacity to commission top-down journey mapping at all. Plenty of practitioners will read those steps and know immediately that a formal exercise would take months to commission, longer to resource, and might never be sanctioned. McKinsey's work on agent deployment is useful here, because it locates the largest opportunity not at the journey level but at the level of individual workflows and tasks. That is precisely where people already know what they do, how long it takes and where the friction lives. They do not need a consultant to tell them. They need a structure that lets them capture it.
Which is why the practical entry point is usually a bottom-up demand map: give individuals and teams a way to map their own tasks, identify where automation applies, and articulate their demand in structured terms. Individual playbooks become team playbooks. Team playbooks surface patterns. Patterns become the demand intelligence that makes the strategic conversation upward possible. The journey map emerges from aggregated evidence rather than preceding it. Bottom-up and top-down are not competing approaches. One is the ideal architecture; the other is the entry point most organisations can actually reach.
The wider evidence is what makes this urgent rather than tidy. The OECD's September 2025 report on governing with artificial intelligence examined 200 AI use cases across governments worldwide and found most national initiatives trapped in the pilot phase. The UK Public Accounts Committee finding quoted within it is the starkest line in any official document I have read: "no systematic mechanism for bringing together learning from pilots, and few successful examples of at-scale adoption across government." McKinsey's 2026 survey of AI trust maturity found only one in three organisations at a meaningful maturity level in agentic AI governance. In public sector contexts, where accountability is non-negotiable and recovery from error is politically costly, that gap is a liability accumulating in real time.
Then automate the intake, deliberately, because that is what protects the part of the job only a person can do. Not to commoditise the role but to rescue it. When you are no longer the intake form, you have the space to read the room in a tense meeting, and to decipher the anxiety behind a partner's demand for a shiny new system.
The scarce good is the authority to pause
Here is the further connection, and it is the one I would put in front of an executive committee.
In an AI-enabled organisation the scarce good is not speed. It is the authority to pause.
The highest value in this work is often deceleration: interrupting a room that has already reached its conclusion to ask whether the problem was ever properly defined. That intervention has always been valuable. What has changed is the cost of not making it. When execution is slow, a misdefined problem produces a disappointing project. When execution is fast and automated, a misdefined problem produces a disappointing project at scale, quickly, with an audit trail that makes it look deliberate.
Note that this is a standing problem rather than a personality one. Anyone can raise an objection. Very few people can stop a room without simply being routed around, and the difference between those two positions is not seniority. It is accumulated trust, and trust is something institutions build or destroy through how they organise themselves.
The choice
Every organisation now faces the same binary. Either someone in the room is asking the hard questions before commitment, or those questions get asked by a post-implementation review when the money is already spent.
The technology is arriving faster than the readiness to receive it, and the gap between the two is not a technical gap. It is a demand gap. Demand is shaped by people, upstream, before the contract.
You do not need a new framework. You need permission to question the axioms, the standing to point at the territory when the map is better designed, and someone whose actual job is closing the loop.
Critical thinking is not a soft skill. It is infrastructure. Treat it accordingly.
References
Gartner (June 2025), Over 40% of Agentic AI Projects Will Be Canceled by End of 2027; and Gartner on agentwashing and the position of AI agents on the Hype Cycle.
MIT NANDA (August 2025), The GenAI Divide: State of AI in Business 2025.
OECD (September 2025), Governing with Artificial Intelligence: The State of Play and Way Forward in Core Government Functions, including the UK Public Accounts Committee finding quoted above.
McKinsey (March 2026), State of AI Trust in 2026: Shifting to the Agentic Era; and McKinsey research on AI financial return at scale and on agent deployment at workflow and task level.
IBM Institute for Business Value, on enterprise AI return and anchoring in business need.
Deloitte, on AI investment payback timelines.
Satya Nadella (June 2026), public remarks on the future of the firm, human capital and token capital.
Gennaro Cuofano (June 2026), Satya Nadella Just Described Harness Theory Without Naming It, FourWeekMBA.
Jean Baudrillard, Simulacra and Simulation (1981), for the concept of the simulacrum.
Kecheng Liu, Professor of Informatics, Henley Business School, for the inverted onion.
BRM Institute, Business Relationship Management Body of Knowledge, on demand shaping, value harvesting, strategic partnering and relationship maturity.
Government Digital Service, AI Playbook for the UK Government; Public Sector AI Adoption Index 2026.
About the author
Sebastian Moore is a Business Relationship Manager working in the UK public sector, at the point where government ambition meets delivery. He was named in Apolitical's Government AI 100 and recognised as a Global Top BRM by the BRM Institute, and he co-founded the Civil Service Care-Experienced Network.
Make sure to share your own thoughts with the author by leaving a comment below
Log in or sign up to continue the conversation