Welcome to the first article in a series on experimentation in the government. This first article will give you a primer on what experimentation is, how it’s done, and the value it provides.

What is an experiment?

What do you picture when you hear the word experimentation? Most probably picture some glass beakers, a Bunsen burner, or some wild haired scientist in a lab coat mixing various coloured liquids together. But the truth is, we are all scientists at heart, and we perform some of the basic principles of experimentation quite often. If you have ever taken a different route to work to see if it was faster, in a sense, you performed your own little experiment. You had a hypothesis that a new route might be faster based on what a colleague told you, you used the same car and went at the same time of day so you could rule out traffic patterns, and you had a way to measure the success of your hypothesis (time to get to work). So, by having an educated guess that a new way might be faster, controlling for outside influence (you used the same car and left at the same time), and having a precise way to measure which way is best, you have utilized the principles of experimentation.

Some people may recall the Pepsi Challenge, a highly publicised ‘blind’ taste test created back in 1975 that set out to scientifically answer the age-old question ‘What tastes better, Pepsi or Coke?’. The experimenters had the right idea in removing some of the bias by not showing participants the labels (the blind aspect), however, as further experiments were conducted, it seemed that the odds were in favor of Pepsi apart from the taste. As it turns out, Pepsi is slightly sweeter than Coke, and people tend to prefer sweeter beverages in small sample sized doses (but whoever has just a sip?). In addition, Pepsi was always labeled with the letter ‘M’, and Coke with the letter ‘Q’, and people generally tend to prefer the letter ‘M’ over ‘Q’. Turns out that when these factors are taken into account, there is no discernible difference in taste. The moral of the story is that a poorly designed experiment can lead to poor decision making... whether you’re deciding which type of drink to buy, or whether your decision affects millions of people.

So how is an experiment done in the government?

People in the government are testing different ideas all the time, but without any training in experimentation, they may not be getting the most accurate data to base decisions on. Similar to the issues with the Pepsi challenge, they may not be considering all the ways in which their testing environment might be skewed, or how their data might be biased. For example, how sure can you be that the data you collected reflects a general trend and is not just a fluke? A coin theoretically when flipped a hundred times should be heads half of the time, but it rarely does, but at which point do we conclude we have a trick coin? This is just another example of the aspects we consider when running experiments that make it more reliable than just simply pitting one idea against another.

As an example, one of our clients were having trouble getting people to click on links in an email that sent them to registration pages for various government programs. So, we did some background research and found that the tone of language within an email can influence people’s decision to perform various tasks such as filling out forms or clicking on links. With this in mind, we crafted four different emails with four different tones of language (e.g., a ‘friendly tone’), and participants in our experiment either got one of those or the original email. We then compared each email to see which one had the highest amount of clicks. And because I know you’re curious, in this case, it turns out that people are more likely to click on a link if language has a friendlier tone!

When we were designing the experiment, the client mentioned that their email distribution system was manual and therefore they could only send one version of the email per day, which meant that one version would go out on a Monday, and another version on a Tuesday etc. This could be a problem. Think about how motivated you feel on different days of the week and how this might affect how you respond. Long story short, we recommended that they extend the testing period and counterbalance the emails (make sure that each version had a chance to appear on each day of the week), therefore equally distributing the motivational qualities of each day of the week. And voila, we created an environment that produced reliable and unbiased data.

Keep in mind, this is just one way we can experiment in the government. We’ve done a myriad of experiments including testing new features of AI chatbots, webpage designs, and safer login procedures. All of which can look completely different, but the same systematic approach and attention to detail still applies.

Why experiment in the government?

Simply put, using unbiased data reduces the risk of costly mistakes. When a pilot lands a plane, they use the data from their instruments to guide them, they don’t just trust their gut. The same goes with making decisions in the government.

Case in point, another client came to us and said that people seem to be making mistakes at a particular point in an application process that they manage, and it was causing an influx of calls coming into call centers. More calls equals more taxpayer dollars and longer wait times, and more frustrated people. After chatting with the client, they mentioned they were going to go ahead and change that part of the application process to something new. Essentially, they asked some people involved in the process and got their opinion on how they should fix the problem. Before launching, we suggested comparing the old way with the new change, just to be sure. Low and behold, after testing with a small group of participants, it appeared that people made less mistakes with the original design than the new one! This goes to show that you shouldn’t always ‘trust your gut’ when making predictions on how people will behave, even if you have experience with that subject. Based on our test data, there was almost twice the amount of mistakes made with the new design, and who knows how many more calls would have been coming in if they went ahead with the change without proper testing.

Final Thoughts

In reflecting on the experimental method in the government, I’m reminded of what a colleague once said, ‘there’s never enough time to do it right, but there’s always enough time to do it over’. I think it’s human nature to not want to pay the upfront cost for something, like purchasing insurance when you rent a car. But ultimately it is up to you, so ask yourself, ‘how safe do you need to be?’ How much is at stake here? How much will it cost to redo it?

About the author

Jesse Howell has a PhD in neuroscience from Carleton University, where he studied how genetic mutations interact with lifestyle factors such as diet and exercise in predicting the risk for mental illness. Currently, he is applying these research skills to service design as an experimentation consultant at Innovation, Science and Economic Development Canada.

He can be reached at: jesse.howell@ised-isde.gc.ca

Image Credit: https://unsplash.com/@alexkondratiev


Make sure to share your own thoughts with the author by leaving a comment below