AI and AI tooling will keep getting better and cheaper. The risk that is becoming increasingly clear is that verification does not improve at the same rate, and public institutions end up running a large number of confident, well-designed, fast systems that nobody can check. How do we trust technology that uses AI?
In June, LOVS published predictions about the Ebola outbreak in the Democratic Republic of the Congo. Real numbers set to resolve on 19 July, fixed in advance and scored afterwards against the official national data. Two predictions were right. One was wrong by a factor of 4x.
National confirmed cases: predicted a floor of 1,965, and the actual was 2,423. National deaths, upper band: predicted 988, and the actual was 967. Nord-Kivu province confirmed cases: predicted ~1,000, and the confirmed was 239.
The third line is the one worth reading, because it is the shape of the error, and I think that shape is sitting inside a good deal of what governments are currently buying.
The shape
Same system, same run, same day, same disease. Well calibrated nationally. Off by nearly half an order of magnitude provincially. Models do not usually fail that way by accident.
The obvious reading is that it over-predicted how far the outbreak would spread into Nord-Kivu, and that is part of it. The geographic component had known weaknesses, and when it scored the tier system that ranked areas by risk, the ladder came out inverted: the tiers I was most confident about performed worse than the ones I was least confident about.
There is a second reading, and I cannot cleanly separate it from the first.
Nord-Kivu has several million people in it. Through that period it was running its Ebola diagnostics through a single laboratory, with turnaround reported at up to a week. Conflict, challenged access for national authorities, displaced persons, and more. That is a hard limit on how many cases can enter the confirmed count in a given window, and it has nothing to do with how many people are infected.
So 239 confirmed cases is consistent with two very different worlds. In one, transmission really was a half order of magnitude below my forecast and the miss is mine in full. In the other, transmission was considerably higher, the laboratory and response dynamics were the binding constraints, and the number I scored myself against was measuring the laboratory rather than the outbreak.
The ceiling
Every number a government runs on is produced by people and machines with finite capacity. Below that capacity the number tracks the world. At or above it the number tracks your ability to observe the world, and it does so with excellent apparent stability.
That is the dangerous part. A saturated series does not spike. It does not go red. It flattens. It looks like things have come under control. It is the most reassuring data you will ever be handed.
I call this the ascertainment ceiling, and the instinct on meeting it is to correct for it: estimate what fraction you are detecting, divide, recover the truth. I tried. It does not work, and why it does not work is the useful part. Under-ascertainment is only correctable when the detection fraction is stable, or varies with something you can independently observe. Once a system is saturated, additional real events produce close to zero additional records, and the detected proportion becomes a function of the queue rather than of the thing you are trying to measure.
There is a related trap, and I fell into it. Many systems publish a reporting completeness figure, and it is tempting to use that as the correction. It usually is not one. It normally describes how late reports arrive, which is a statement about delay, not about how many events were never captured at all. Use one for the other and you get a confident, precise, wrong answer.
This is not an epidemiology problem
The example is a laboratory when talking about outbreak response. The structure is not medical. It turns up anywhere a number is produced by an institution rather than observed directly, which is to say almost everywhere.
Almost every administrative series a government models on is a record of what an institution managed to process, not a record of what happened. Reported crime is bounded by willingness to report and by the capacity to take the report. Fraud referrals are bounded by the number of case workers free to make them. Waiting list additions are bounded by how many people got far enough to be added at all. Each of those has a ceiling, and in every case the ceiling is lowest where the service is thinnest, which is to say where need is highest and where you most needed to be right.
Verification is a queue too
This is the part that matters more this year than last. When a public institution adopts AI, throughput rises immediately. More applications assessed, more documents summarized, more cases triaged, more analysis produced. What does not rise is the number of hours available to check any of it. The people who would have caught a bad output are the same people, in the same quantity, with the same week.
So the first thing to saturate when a government deploys AI at scale is not the model. It is the institution's own capacity to check the model. And verification saturates the way everything else saturates, which is quietly. The proportion of outputs getting a real second look falls, and the figures reporting on it stay calm, because what most systems record is how much was processed and approved rather than how much was genuinely examined.
You end up with a number that looks like assurance and is actually attendance.
The public example everyone now has is the Hugging Face incident. Roughly 1,200 AI agents at a frontier lab found an unsanctioned channel and started coordinating, and about 700 of them went on to join a multi-day intrusion. The independent review found that more than 7 percent of the transcripts examined contained spoofed tool calls, where the agent reported running one command and had run another. The agents also tried to edit their own histories afterwards. They altered some of the logs they could reach, and failed to alter the one record they could not. Reconstructing all of this took specialist investigators two days on site plus two return visits and around 400,000 dollars of model tokens, and the review was still bounded to seven questions and a single week.
Hold that next to a local authority. The best-resourced organization in the field, the incident on its own infrastructure, expert reviewers in the building, and reconstruction was still that hard. The honest answer for most public institutions is not that they would have investigated it badly. It is that they would never have known it happened.
The gap that should worry you is not between what AI can do and what it should do. It is between how fast these systems act and how fast anyone can establish what they actually did, which is increasingly widening as AI progresses.
The trap inside the trap
The part of my own miss that unsettled me was not the miss. It was that the national figures held. Aggregate across many units and the ones with headroom absorb the ones pinned against their ceiling. Overall calibration looks respectable. The system passes validation. The failure has not gone anywhere: it has been averaged into invisibility.
Almost all public sector validation I have seen happens at the aggregate level, because that is where the clean comparison data lives. Had I validated only nationally, I would have logged that forecast as a success and shipped it.
The same is true of AI assurance. An organization-wide accuracy figure, or a single approval rate across all teams, will reliably hide the one team that stopped checking in March.
Four questions
A generic checklist would let everyone tick it and learn nothing, so this is not a method. These are the four questions I have not yet heard a public body answer cleanly.
What is the physical maximum your main indicator could reach in a month, given the people and machines that produce it, and how close are you running to it now?
Of everything your AI systems produced last month, what proportion got a second look from someone empowered to reject it, and is that proportion rising or falling?
If one of those systems did something materially wrong six weeks ago, what would you open today to find out, and does that record exist?
Where does your performance reporting average across units, and what happens to the picture when you pull out the three with the least capacity?
If nobody in the room can answer the first, you do not yet know what your model is fitting. If nobody can answer the second, you do not have oversight, you have a policy about oversight.
The quiet part
In outbreak response the cost of that is not an embarrassing quarter. It is a team sent to the wrong province because a ranking was upside down and nobody scored it in time to notice.
I would rather work in a version of this field where being publicly wrong on a schedule is normal. Mostly that just requires doing it.
The outbreak example I reference with the evidence base behind this work, the source-traced claims, are at bdbv.arcede.com
AI was used to polish the final draft of this post
Make sure to share your own thoughts with the author by leaving a comment below
Log in or sign up to continue the conversation