Across governments, a quieter transformation is already underway. Civil servants are informally adopting generative AI tools to manage growing workloads, creating what can be understood as a “shadow AI” environment operating outside formal oversight. While this shift offers significant productivity gains, it also introduces new institutional risks, including reduced visibility, blurred accountability, and subtle changes in how decisions are made. The challenge is no longer whether AI will be adopted, but how governments respond when adoption has already begun.
Governments around the world are moving quickly to define how artificial intelligence should be used in the public sector. Strategies are being published, guidance is being updated, and regulatory frameworks are evolving to manage risks around privacy, bias, and accountability. The underlying assumption is clear: AI adoption is something institutions can plan, control, and roll out in an orderly way.
But generative AI is not entering government in that way.
Across organisations, including the public sector, employees are already using tools like ChatGPT, Copilot, Claude and Gemini in their everyday work for drafting documents, summarising reports, and analysing information. Often, this happens informally, outside formal approval structures, and sometimes without the organisation fully knowing. Recent research describes this as “shadow user innovation” referring to the covert use of generative AI when the benefits of using it outweigh the perceived risks of disclosure.
This creates a governance mismatch: AI is entering through practice, while policy is trying to control it through design.
The scale of this shift is already significant. The 2024 Work Trend Index from Microsoft and LinkedIn finds that 75% of knowledge workers use AI, and 78% are bringing their own tools into the workplace. At the same time, many employees remain hesitant to disclose this use, creating a visibility gap. Shadow AI is not just about adoption; it reflects the fact that organisations often cannot fully see where, how, and why these tools are already shaping work.
In government, this gap is particularly consequential. Public institutions are not only concerned with efficiency, but with accountability, fairness, and trust. When AI use becomes informal and invisible, those institutional guarantees become harder to uphold. But even when AI use becomes visible, a deeper problem remains.
Governing the wrong problem?
Much of the current policy response treats AI risk primarily as a question of accuracy and compliance.
For example, UK government guidance encourages civil servants to use generative AI cautiously: avoid entering sensitive data, treat outputs as potentially misleading, and verify results before use. At the same time, it recognises legitimate use cases, including research support, summarising public information, and secure analytical applications.
That practical emphasis matters. It shows that even relatively cautious frameworks are already moving beyond the binary of “allow” or “ban.” But the dominant logic is still defensive: protect data, reduce misuse, check outputs.
Similarly, broader governance discussions emphasise risks such as hallucinations, bias, and data misuse. Recent OECD Analysis warns how AI systems can produce convincing but incorrect outputs, that automation bias may encourage uncritical reliance on machine-generated recommendations and that overreliance may weaken independent judgement.
At the same time, governments are already experimenting in practice. The OECD points to an Australian Public Service trial of Microsoft Copilot across more than 60 agencies and over 7,700 public servants, where users reported gains in efficiency and drafting capability, even as capability and cultural barriers remained. It also highlights Spain’s GovTechLab, an AI use case incubator, which identified around 300 public-sector generative AI use cases and selected several for real-world pilots in areas such as document classification, AI assistants, and tenders and grants.
These examples are useful because they show that the issue is no longer whether governments are experimenting with AI. They are. The question is what kind of governance logic is emerging around that experimentation. Taken together, current approaches reflect a shared assumption that improving accuracy and controlling inputs will be sufficient to manage risk.
However, this assumption is increasingly being challenged.
Rethinking Risk: When AI Shapes Judgement
A recent paper, “Beyond Accuracy: Rethinking Hallucination and Regulatory Response in Generative AI,” argues that focusing on accuracy alone can obscure the real nature of harm. As AI systems become more fluent and statistically reliable, users may begin to trust them more, even when they are incomplete or misleading. This creates an “accuracy paradox” as improvements in accuracy can increase overconfidence and reduce scrutiny.
In other words, the risk is not only that AI produces incorrect answers. It is that it produces answers that feel authoritative enough to go unquestioned.
For public administration, this introduces a deeper layer of institutional risk.
First, it creates a form of epistemic instability where persuasive but flawed outputs can quietly shape policy advice, administrative reasoning, and evidence interpretation. These distortions are difficult to detect precisely because they appear credible.
Second, it generates a responsibility gap. While official guidance is clear that civil servants remain accountable for decisions, the informal or undisclosed use of AI makes it harder to reconstruct how those decisions were formed. When AI becomes part of the reasoning process but remains outside formal systems, accountability becomes blurred.
And third, it raises a more fundamental concern. Generative AI systems do not reason in the way bureaucratic institutions are designed to. They optimise for plausibility, not truth. If administrative judgement becomes increasingly mediated by such systems, the very nature of the “public service mind”, grounded in deliberation, justification, and traceability may begin to shift.
These are not just technical risks. They are institutional ones.
Seen in this light, the challenge of shadow AI becomes more complex. It is not just that employees are using AI informally. It is that they may rely on it in ways that are difficult to observe, trust it in ways that are difficult to regulate, and normalise it in ways that may scale across entire organisations. These risks do not disappear as systems improve. In some cases, they intensify.
This raises an important question for public administration: are governments governing AI systems, or are they trying to govern the wrong problem? If policy focuses primarily on accuracy and compliance, it may miss how AI is already reshaping judgement, workflow, and discretion from within.
Learning from practice: Boston’s “middle path”
One of the more interesting responses to this challenge can be seen in the City of Boston.
Rather than framing AI as something to prohibit or fully embrace, Boston has taken what is often described as a “middle path”, recognising that employees are already using these tools, and focusing on how to build governance around that reality.
This approach begins with visibility. Internal surveys found that around 6 in 10 employees had tried generative AI, and about 1 in 4 were using it regularly at work. Workers expressed strong interest in training, and the most common use cases were practical ones: writing, summarising, and data analysis. At the same time, the leading concerns were inaccuracy, security, plagiarism and intellectual property, and data privacy.
Instead of suppressing this activity, Boston has sought to channel it through institutional design. This includes providing secure, city-sanctioned environments for AI use, reducing the need for employees to rely on personal tools. It involves testing systems in low-risk contexts such as open data, where experimentation does not compromise sensitive information. And it embeds human-in-the-loop oversight structures, ensuring that AI outputs remain subject to human judgement rather than replacing it.
In one example, the city introduced an AI-powered search tool on Boston.gov to help residents navigate public information more easily. According to Boston’s published update, as of December 2025 the positive feedback for AI search reached 34.3%, compared with 10.9% for traditional search, representing a difference of 23.4 percentage points.
This is not simply an adoption story. It is a governance approach grounded in practice rather than assumption. What distinguishes this approach is not the technology itself, but the logic behind it.
Instead of asking, “How do we prevent misuse?”, it asks, “How do we understand and shape how this is already being used?”
What Boston demonstrates is not just a local solution, but a broader shift in how AI is beginning to be governed in practice.
Taken together, these examples and emerging research suggest that the core challenge is not just that AI is being used informally, but that governments are governing it through an incomplete lens. Shadow AI reveals a visibility problem. The accuracy paradox reveals a trust problem. Behavioural risks reveal a decision-making problem. Each of these points to a different dimension of governance. But current approaches often collapse them into one: compliance.
This matters because AI is not simply another administrative tool. It is beginning to influence how information is interpreted, how options are framed, and how decisions are made. In that sense, the governance challenge is no longer only about controlling systems. It is about understanding how those systems reshape human judgement within institutions.
For governments, this suggests a subtle but important shift in emphasis. Not only regulating use, minimising risk, and ensuring compliance, but also making use visible, creating safe channels for experimentation, and recognising how AI changes the conditions under which decisions are made.
Because if generative AI is already embedded in everyday work, governance cannot begin at adoption. It has to begin at use.
Shadow AI, in this sense, is not just a problem to solve. It is a signal. It reveals where institutional processes are under pressure, where tools are already filling gaps, and where governance models may be lagging behind reality.
The question is not whether this signal exists, but what governments choose to do with it. More fundamentally, if AI is already shaping how decisions are made, the real risk is not that governments lose control of the technology, but that they misunderstand the nature of control itself.
Log in or sign up to continue the conversation