Search across all content
A proof-of-concept trial testing whether a large language model could produce a workable first draft of a clause-by-clause explanatory note for amendment Bills, for a drafter to review.
The Parliamentary Counsel Office is the office that drafts New Zealand's laws and publishes them for the public. One part of that job is writing explanatory notes, the brief summaries attached to a Bill that tell readers in plainer terms what each numbered provision is intended to achieve. They help legislators and the public understand what a law is intended to do.
Writing them is demanding, especially for amendment Bills, which change an existing law. To explain an amendment clearly, a drafter has to work through the proposed change against the existing law it alters, known as the principal Act, and spell out the effect on it, provision by provision. The work is detailed and exacting; it is sometimes done against the clock, and the Office carries this work alongside a growing drafting workload.
As part of a six-month research and development programme that looked at several possible uses of AI across its work, the Office worked with the technology company, Catalyst, to test one narrow thing: whether a large language model, a type of AI trained on large amounts of text to generate language, could produce a workable first draft of an explanatory note that a drafter could then review and build on. The work was deliberately limited to amendment Bills. Explanatory notes are written for many kinds of legislation, and the Office kept the trial to this one type because the project was a proof of concept, where a narrow, well-defined scope kept the experiment manageable.
The prototype takes an amendment Bill, reads it against its principal Act, and produces a clause-by-clause summary of what each change would do. A user can ask for a summary again, adjust it, and then pull it out to work on further. The design kept the model's role narrow. The job of parsing the documents was given to ordinary software code, drawing on the structured, machine-readable form in which legislation is already stored, and the model was used only to write the summaries, steered by carefully designed prompts. The team also preferred models that could be hosted within New Zealand, for data protection and security. Catalyst, the delivery partner, reports that this involved fine-tuning a smaller model, training a more compact language model on a set of New Zealand legislation examples, hosting it on a New Zealand sovereign cloud, and adding safeguards so staff could weigh several options and exercise their own judgement. The things the team set out to test were whether a model could produce a draft good enough for a person to review and work from, what model size suited the task, and whether a smaller model could be fine-tuned to do it well.
1. Feeding the tool smaller sections of text produced better drafts
The team found that the summaries were far more accurate and useful when the model worked on small sections at a time. Handing it a whole Bill to break down, clause by clause, yielded weaker results.
2. The drafter's judgement remained the deciding factor
Because users could generate a draft and then revise it, the exercise showed how much rested on the drafter's decision, based on their own expertise, about whether what came back was worth building on.
3. Model size sets up a trade-off that the Office still has to weigh
Bigger models generally drafted better, but they were more expensive to run, and their mistakes were harder to catch, which left an open question about the standard of draft the Office actually needs. No time savings or other outcomes were measured, since the work stopped at a working prototype. The Office's deputy chief executive for digital services, Andy Neale, has described this as an ambition, suggesting that halving the time it takes to write an explanatory note would count as a success.
The tool was built to assist a drafter, with a person always in review. The Office approached the work at a cautious pace, and has been careful to give the technology only the access it needs, to head off unintended side effects. Nothing in the design lets the model make a decision or produce a finished note on its own.
The conditions around the model mattered as much as the model. The results rested on the legislation already being held in a structured, machine-readable form and on the task being broken into small parts. A government hoping to try something similar would need its laws stored that way, and drafting expertise close at hand to judge what the tool produces.
The main choices were left to formal evaluation. What size of model to use, the quality to aim for, where such a tool would sit in the production process, and how to bring specialists into the shaping of the prompts were all set down as questions for the next phase, left open by the prototype.
The Office shared its findings and code openly. It published a report and released the prototype's open-source code for others working on legislative technology to learn from and reuse. The part of this worth taking elsewhere is that approach: a contained, human-reviewed proof of concept, built on structured data, tested against specific questions, and shared openly before any decision to adopt.
This case study was written with assistance from artificial intelligence.





Connect with 500,000+ public servants solving your hardest challenges.





Connect with 500,000+ public servants solving your hardest challenges.
Help public servants worldwide learn from your work, what worked, what flopped and what you'd do differently
Share your project
Log in or sign up to continue the conversation