This article is written by Jonathan Breckon, Director, and Anna Hopkins, research associate at the Alliance for Useful Evidence.


What is experimentation in policy or practice? And why is an experiment worth investing in? What experimental tools are out there? And when or where might they fall short?

Across the world, a small but growing movement of policy experimenters are bringing fresh ideas to how to solve public problems. From crafting better services, to making the back-office of government more efficient, new methods and tools are being used to develop and test policy.

An Inventory of different designs published today by The Alliance for Useful Evidence and Nesta catalogues the contents of the policy experimenter’s toolbox. In it, we present 18 experimental approaches to evidence-informed policy, and provide jargon-free advice on the pros and cons of each one.

Why experiment?

From the thousands of experiments conducted by Thomas Edison to create the first lightbulb, through to trials in medicine, or the long-running field experiments than underpin modern agriculture, trying ideas out in practice is a cornerstone of scientific and technological discovery.

Experiments are now critical to sectors where innovation and optimisation are routine, like web development. It has caught on in business. Among the largest financial institutions, retailers and restaurants in the US, at least a third are running randomized experiments, with companies like Google and Amazon running tens of thousands of experiments a year. A/B testing is now the standard means through which Silicon Valley improves its online products. However, in government experimentation remains relatively rare.

The word “experimental” has come to mean “innovative” or “radical” rather than simply “untested”. But genuine experimentation is about committing to rigorous evaluation and evidence — not just freewheeling “trying stuff out”

When governments do experiment, the language can be confusing. For some, experimentation is just a synonym for innovation — trying out a new idea, like the introduction of gender quotas in Rwanda, or Citizens’ Assemblies in Ireland. The word “experimental” has come to mean “innovative” or “radical” rather than simply “untested”. But genuine experimentation is about committing to rigorous evaluation and evidence — not just freewheeling “trying stuff out”. It’s about putting in place a structure to learn from trying things out in the world. It is not about just doing things differently — and expecting to succeed.

Experimentation necessitates a more mature attitude from leaders and decision-makers, one that allows us to learn positively from “good failure”. Rather than pretending we have all the answers, it means admitting we that we don’t, and that we need to put our ideas to the test. A more transparent, “open-by-default” approach is being pioneered internationally, by experimenters in Canada, Finland and the US.

The experimental toolbox

At the Alliance for Useful Evidence, which is funded by the National Lottery Community Fund, the Economic and Social Research Council (ESRC), and Nesta, we have made the case that government must rigorously and systematically put policy to the test – or risk stagnation.

Our new inventory catalogues 11 kinds of randomised experiments used to inform decision-making. Of these randomised experiments are perhaps one of the most powerful.

Randomised experiments aim to test a policy idea or innovation by investigating what difference it has made for the people it is aiming to help. They do this by using a control group; testing an innovation against “business-as-usual”. This doesn’t have to be a large trial like testing a drug. They can be fast and flexible. For example, the World Bank are advocates for “nimble RCTs”. They funded nimble evaluations on how best to improve the take up of health insurance in Azerbaijan, expand the use of contraceptives in Burundi, and support teachers to deliver tailored education to children affected by war and displacement in Lebanon.

There are opportunities to experiment even on difficult social issues. But in these cases, creating a control group carries its own risk. If, for example, you are experimenting with a new benefits scheme, having a control group allows you to test whether the scheme has actually achieved its goals, but it also means depriving a group of citizens of an improved service, at least initially.

It’s important to be clear what different experimental designs can and can’t tell us

To get around this, a phased policy roll-out can create what’s called a “waiting list” experiment where the control group are the soon-to-have-innovation group. Everybody eventually receives the new policy innovation, negating the risk of push-back from the public. This approach was used by the UK’s Ministry of Housing, Communities and Local Government in their first ever randomised trial on community integration. In 2016, they tested whether English language training could help immigrants engage more in their community.

The results were impressive. They found significant improvements on many social integration outcomes, such as new friendships formed with people from other cultures, and attending more health appointments. The trial helped the Department put together new plans for their 2018 Integrated Communities Green Paper — including a network of conversation clubs and a new English language fund. For others wanting to design their own experiments in the UK Government, there is a Trials Advice Panel run by the Cabinet Office What Works Team who can walk you through the technicalities.

There are also other valuable tools to help guide public decision-making. Despite their unfashionable status among some policy wonks and evaluators, we also highlight the value and utility of Quasi-Experimental Designs. We explain these methods, which are too often couched in deadly technical language in plain English. These methods have been used to evaluate policies like the Troubled Families programme or the impact of the Sure Start programme of children’s centres in the UK, including newer approaches like Synthetic Control, that use inventive statistical methods to take advantage of the advent of “big data”.

Understanding the difference

It’s important to be clear what different experimental designs can and can’t tell us. Only certain experimental designs are helpful for learning about impact and effectiveness — ie. whether or not our new idea is really making a difference. Here, randomised experiments, or quasi-experimental designs are uniquely valuable. But other approaches may be better suited to finding out different things, at different stages of developing a policy solution.

Prototyping, for example, emphasises front-loading risk and creating a solution with a better chance of success through stakeholder engagement. Using plastic Lego bricks to build prototypes of engineering products is a low-fi, low-resource way of making early operational or design issues obvious. But it won’t tell you whether or not the new system works in real life, or at scale.

As well as increasing our understanding of what experimental design to use when, there is scope for experimenters to get much better at using different tools in combination to innovate more effectively. Prototypers, for instance, could use low-cost randomised trials like A/B tests or nimble trials to evaluate prototyped products or services.

We hope this resource will be useful in clarifying when experiments are useful and why, as well as simplifying jargon and double-speak, and myth-busting some common misperceptions. As innovators seeking social good, we have a duty to put our ideas to the test — to find out what doesn’t work, discover what does, and improve people’s lives. — Jonathan Breckon and Anna Hopkins

Read the Experimenter’s Inventory her****e.

(Picture credit: Death to the stock photo)