Valuable datasets remain buried or locked in formats and systems that render them unusable, so making this data legible is a direct way to create value. Just mapping the invisible city - the hundreds of miles of covered cables, pipes, and infrastructure laying beneath the urban surface - can save millions of dollars. The UK government estimates that its National Underground Asset Register could generate more than £400 million in economic benefits each year by reducing accidental strikes on underground infrastructure.
But making this data usable is expensive. Some governments lack the resources to clean, process and prepare this data for dissemination. If they have to publish it openly and cannot recover the costs, this work can appear as another unfunded item in an already tight budget. An item that never makes it into the priority list, so rich data remains locked in closed systems, when not in physical archives.
This is the issue that the UK’s proposal on charging for certain public-sector data reuse is trying to tackle.
And it’s not merely a technical issue. Decisions on how to fund open data have deep implications on who pays, who benefits, and what are the terms of access to a critical infrastructure on top of which thousands of products and services - public and private - are built.
The marginal cost restriction on public sector data re-use
On July 15th, the Department for Science, Innovation, and Technology (DSIT) launched a call for evidence on the proposal “to create greater flexibility to charge above marginal costs for more public sector data assets without undermining the benefits of Open Data.” The consultation closes tomorrow, September 8th.
Under the current rules, many public bodies may charge only the marginal cost of reproducing, providing and disseminating data. For digital information, that additional cost is often negligible. In practice, this means that much digital data must be made available for free or at a very low cost. As a result, certain data assets are not being published in a valuable form - or may eventually stop being published - because the governments holding it lack sufficient funding. The idea is that giving more flexibility to charge may help them address this issue.
There are already important exceptions to the marginal cost rule in the UK. Information produced outside a public body’s core public task may fall outside the reuse rules. Higher charges are also permitted for certain bodies or datasets required to generate revenue (e.g., HM Land Registry). Libraries, museums and archives have a further exemption.
These are meaningful exemptions already, so which additional datasets does the government envisage public bodies charging for? Which data assets could emerge with new funding? And which existing data and associated services are genuinely at risk because of insufficient resources? The call gives an example of the latter: data on the condition, extent and location of England’s natural assets collected and published by the Natural Capital and Ecosystem Assessment program. The funding for this data will run out in 2029, and charging for its use may be a way to sustain it beyond that point.
From the available documentation, however, the scale and importance of the problem remain unclear. The call for evidence does not provide estimates on how many data assets are funding-constrained, how much investment is missing or which additional datasets might be generated if their holders could charge for reuse.
Of course, those numbers are difficult to calculate, and gathering evidence is precisely the purpose of the consultation. Yet, knowing how many data assets could be surfaced or are in danger, their cost, and their potential value, is critical to ensure that the resulting policy proposals are proportionate to their stated goals. The government’s Data Valuation Framework also recognises that the value of data is not limited to potential revenue: it includes the wider economic and social value created through reuse.
That distinction is critical. Any resulting policy should examine the demonstrated funding gap, the improvement that charging would finance, the uses that higher prices might discourage and the alternatives available. The goal is to balance the incentives and financial sustainability of producing valuable data against the social and economic value lost by restricting access.
Poor governance design could create fragmented, expensive, and unequal access
I do not believe that all public-sector data made available for reuse must always be delivered for free. Collecting, processing, updating and making data openly available comes with a cost for public entities, and we should not assume that general taxation will always provide sufficient or stable funding for such costs.
There is an important distinction on the sources of value and cost that I think are helpful in informing governance decisions on the funding of open data.
First, there are the fixed costs of collecting, cleaning, documenting and maintaining a dataset. Second, there are the marginal costs created by delivering it to a particular user, including computing, bandwidth and support. Third, there is the commercial value that a user may be able to extract from the information.
These categories demand different policy responses. A funding gap in producing or maintaining a valuable dataset may justify a new financing arrangement, potentially including charges above marginal cost. Disproportionate infrastructure use may justify rate limits, service tiers or usage charges reflecting the additional cost imposed. A user’s willingness to pay may demonstrate demand, but it is not, on its own, evidence that charging will create greater public value.
AI agents make the second issue especially visible. An agent helping a business identify procurement opportunities, a resident understand a planning restriction or a city assess building-retrofit options may need to make repeated calls to several official sources while completing a single task. Mature agents will increasingly query APIs, databases and services rather than merely crawl webpages. That use can generate real infrastructure costs, but it does not automatically justify charging above marginal cost for the underlying data.
So, in principle I am in favor of exploring additional ways to cover the costs of generating and disseminating certain data. But I also think that details matter - a lot.
Here are four aspects of the UK’s proposal that I think are critical to get right:
First, charging should have a clear and demonstrable relationship to producing more data, improving its quality or sustaining the infrastructure required to provide it. Allowing public bodies to charge simply to “improve public services” is too broad and could quickly lead to scope creep. Revenue should normally be ring-fenced for producing, maintaining and improving public data and its infrastructure.
Second, the reform should establish clear criteria for determining which datasets may be subject to charges. Selection should be supported by evidence of the funding gap, the costs preventing publication or improvement, the expected additional value from investment, the users likely to be priced out and the viability of alternatives such as public funding.
Third, fee design should address both access and use. Charges should not create a data playing field tilted towards organisations with deeper pockets. At the same time, unusually intensive users should bear the marginal costs they impose on public-data infrastructure. Keeping these issues separate would allow a core dataset to remain broadly accessible while higher-volume or enhanced delivery services recover the additional costs they create. The example of Wikipedia Enterprise, which leaves Wikipedia’s underlying content free and open but charges commercial organisations for high-volume, high-frequency and real-time access through infrastructure designed for their needs, is an interesting example that governments could explore.
Finally, the rules governing exemptions should be clear, specific and subject to consistent oversight. Leaving each public body to devise its own criteria, terminology, licence and pricing model could create confusion, fragmentation and pockets of abuse. Central rules and supervision do not require every operational decision to be made in Whitehall, but they should give users a predictable system across government.
These are some of the questions that the UK government is seeking input on. Despite my criticism that the consultation lacks concrete estimates of the scale of the problem and the datasets it might unlock, DSIT is trying to fill part of that evidence gap with this call. The initiative deserves serious engagement.
The consultation remains open until 11:59 p.m. on September 8th.
How we fund public data will impact how value is created and who benefits from it. This is a critical piece in the emerging infrastructure of the AI economy, so if you have thoughts to contribute, chime in!
***
This post was also posted in the Datapolis substack.
Make sure to share your own thoughts with the author by leaving a comment below
Log in or sign up to continue the conversation