Search across all content
The State of Wisconsin’s Department of Revenue (DOR) won a Federation of Tax Administrators (FTA) award in 2025 for a project intended to “modernize outdated computer systems and process tax documents and inbound mail systems,” according to an FTA release.
Wisconsin used generative AI-based image recognition technology to securely get to a higher state of accuracy on scanned tax forms. This includes not just the income tax, but any kind of document scanned by the state’s revenue department.
Wisconsin spent February through June converting all inbound mail. By the end of June, the state was able to switch over to process three to four hundred thousand paper items. The department plans to expand the system within the next two years to handle over seven million pages of data being captured annually.
Before using AI in this way, Wisconsin followed standard practice in capturing data from papers. High resolution images were run through software, which did traditional optical character recognition. That process had a relatively high error rate for typewritten documents (different fonts, for example, could easily confuse the software). Additionally, by the time the documents went through various stages, their clarity was reduced. “It’s kind of a photocopy of a photocopy problem,” said Ryan Minnick, FTA chief operating officer.
And it was somewhat less accurate for handwritten documents, in which, for example, it was easy to confuse the number 7 with the number 1.
The new AI-based technology is paying off in several ways. It showed a 300 percent per hour increase in document processing and a 56 percent savings in labor costs, which could allow overworked staff members to focus on other necessary tasks. according to Keith Gross, section chief of Wisconsin DOR's Division of Technology Services.
The model has been trained to be very accurate. Initially, people are still verifying it, but as the model progresses, it improves. Any kind of generative AI-based image recognition tool is not only returning the data, but it’s also indicating how confident it is based on all the parameters that are set. If the tool says, ‘I’m 98 percent sure I got this right,’ and the user agrees that it’s 98 percent accurate or 99 percent accurate, the user can start to rely on the automation a little bit more. Instead of sending every single image to a person to validate, fewer go for manual review.
Although the system has had a great deal of success in producing accurate outcomes, the Department of Revenue recognizes that there needs to be ample human intervention. In fact, said Wisconsin’s Gross, “in the early days of using the system, we frequently found cases where the AI-captured value was correct and the verifier altered it to an incorrect value. So, even though everybody says human is the gold standard, we always knew that humans made mistakes too.”
With that in mind, Gross said, “What we’ll probably be doing in the near future is another study to ask, ‘Are we reaching the point where a correction from a verifier is more often making the data more or less correct?’ And if they make it incorrect more often that might be the point where we realize we need to cut back human review and make them only look at the stuff that is clearly questionable based on rules. I don’t know that we’re there yet. The early ones we felt were heavily influenced by people learning and adapting to the system.”
Attribution: Case study provided to the AI Navigator courtesy of the IBM Centre for The Business of Government, Apolitical's content partner. “AI in State Government: Balancing Innovation, Efficiency, and Risk”, https://www.businessofgovernment.org/reports/ai-in-state-government, IBM Centre for The Business of Government.
In partnership with
IBM Center for The Business of Government
In partnership with
IBM Center for The Business of Government





Connect with 500,000+ public servants solving your hardest challenges.





Connect with 500,000+ public servants solving your hardest challenges.
Help public servants worldwide learn from your work, what worked, what flopped and what you'd do differently
Share your project
Log in or sign up to continue the conversation