Search across all content
A multi-agent AI system, run entirely on the agency's own infrastructure, that turns free-text IT requests into structured tasks.
The Federal Employment Agency (Bundesagentur für Arbeit, or BA) is one of Germany's largest public institutions, managing employment services and welfare processes for millions of people across Germany. One of its major IT systems, known as ALLEGRO, supports the delivery of social benefits, is used by more than 40,000 staff, and handles transactions worth several billion euros each year.
When something in ALLEGRO needs to change, the request is written up in a formal document. These documents vary widely, from a few sentences to twenty pages, and describe the problem, the system affected, and the desired outcome. Before any work can begin, each document needs to be turned into a structured task in Jira, the platform the agency's IT teams use to plan, assign, and track their work. That meant staff reading through each document, pulling out the relevant details, and manually creating a task entry with the right title, description, category, and priority. With around 1,000 of these requests arriving every month, it was repetitive work that consumed time better spent on analysis and oversight.
The challenge is set to grow. The BA expects around 40,000 of its employees to retire or leave by 2032. Any solution also had to meet strict data protection requirements, with no data leaving the agency's own infrastructure.
The BA partnered with Capgemini, a technology consultancy, to build an AI-powered system that could handle the conversion automatically. Rather than relying on a single AI tool, they designed a system in which four separate AI components, known as agents, each take on a different stage of the work.
The first reads through the original document and identifies the key information: what the problem is, which system is affected, and what needs to happen. The second organises that information into a structured set of steps. The third drafts a complete task entry in Jira, with the right title, description, category, and priority. A fourth agent, designed to check each draft for errors, inconsistencies, and duplicate entries, is currently being developed and will be added in a future version.
Nothing enters the system without a person reviewing and approving it first. Every task entry the AI produces is checked by a member of staff before it goes live. All outputs are logged, so there is a clear record of what was generated and what was approved.
The entire system runs inside the BA's own infrastructure, with no data leaving the agency. The AI models it uses, including Aleph Alpha, LLaMA, and Mistral, are open-source or privacy-compliant models that can be hosted internally rather than relying on external cloud services. These are coordinated using CrewAI, an open-source platform for managing systems where multiple AI components work together. Because some of the original documents run to twenty pages or more, the team also built in processes to handle long documents and ensure the AI could work within the limits of what the models can process at one time.
The technology was introduced alongside training and support for the teams using it, so staff understood how the new workflow operated and could maintain quality as it became part of daily operations.
1. In over 80% of test cases, the system produced task entries that domain experts judged ready to use
The research team evaluated the system using 50 authentic documents drawn from across the agency's IT operations, covering a range of lengths, formats, and levels of complexity. Three experienced staff members, each with an average of twelve years in the field, independently assessed whether the task entries the AI produced were accurate, complete, and usable. In more than 80% of cases, they were. The evaluators largely agreed in their assessments, with a high level of consistency across their independent reviews. Early testers within the agency described the outputs as “impressively accurate” and “immediately usable.”
2. The agency estimates the system saves 100 to 130 working hours each month
The research found that each task entry took an average of six to eight minutes less to produce than when done by hand. With around 1,000 formal change requests arriving each month, that amounts to roughly 100 to 130 hours of staff time freed up, time that can go toward analysis, decision-making, and oversight rather than extracting and reformatting information.
3. Task entries became more consistent across teams
When multiple people create task entries manually over time, differences in structure, language, and level of detail tend to accumulate. Because the AI applies the same approach to every document, the entries it produces follow a consistent format regardless of who submitted the original request or which team is involved.
1. Running AI entirely within government infrastructure is achievable, not just theoretical. The BA's data protection requirements meant no information could leave the agency's systems. By using open-source and privacy-compliant AI models hosted on its own hardware, the agency demonstrated that a multi-agent AI system can operate within a secure government environment without relying on external cloud services. The research paper describes this as a proof point for other public institutions operating under similar constraints.
2. Introducing the technology without supporting the people using it would not have been enough. Training and support for staff were built into the rollout from the start. The research team observed that technology on its own does not drive lasting change and that staff needed to understand the new workflow to maintain quality and trust in the system's outputs.
3. The system struggles with long or ambiguous documents, and the research team was transparent about that. Long documents, those exceeding the amount of text the AI models can process at once, sometimes produced incomplete outputs in roughly 8% of cases. Documents that were internally contradictory or ambiguous caused problems in about 5% of cases. In those situations, the time spent correcting the AI's output could exceed the time saved, and the system functioned more as a drafting aid than an efficiency gain. The feature for estimating how long a piece of work will take is not yet operational, and the fourth agent, designed to check for errors and duplicates before human review, is still being developed.
This case study was written with assistance from artificial intelligence.





Connect with 500,000+ public servants solving your hardest challenges.





Connect with 500,000+ public servants solving your hardest challenges.
Help public servants worldwide learn from your work, what worked, what flopped and what you'd do differently
Share your project
Log in or sign up to continue the conversation