This article is Part 9 of Sally Washington's Building Policy Capability series, which looks at how to improve the work of government through high-performing policy systems. 'Looping the loop' explores how to embed evaluation throughout policy processes.


Decades ago, I lamented the dearth of evaluation in Aotearoa New Zealand’s public policy system. My paper ‘Looping the loop’ ended up informing a co-authored chapter in a World Bank ‘Atlas of Evaluation’, which showed that Aotearoa was not alone in this gap in its policy advisory system. The situation hadn’t changed much when in 2014 I diagnosed Aotearoa’s policy capability while setting up the Policy Project in the Department of the Prime Minister and Cabinet. And it has been a common theme in the range of jurisdictions I’ve worked with since.

In 2022, OECD countries adopted the ‘Recommendation of the Council on Public Policy Evaluation’ committing member countries to:

  • Promote the quality of evaluations, establish standards, and develop skills and capacities,
  • Conduct evaluations, communicate the results, and embed them into decision making processes.

More evaluation activity isn’t enough. For evaluation to be “embedded in decision making” it needs to feature in upstream policy advisory processes. All too often ex-post evaluations are used for accountability purposes as opposed to generating insights about what worked, or didn’t, and how that information can help improve future policy processes. This includes not only what worked in a substantive sense, but what worked in terms of process, methods, choice of policy instruments and ways of working. Evaluation should not be just about ex-post review. Indeed, traditional forms of ex-post evaluation and the project management systems that underpin them can sometimes work against agility, adaptation, learning and continuous improvement along the way.

In increasingly complex environments where governments are not the only actors in delivering results for people and communities, there are methodological challenges with assessing contribution to outcomes (when multiple actions and actors combine to make something happen). Timeframes also need to be reconsidered. Government demands to demonstrate ‘quick wins’ are incongruous in areas where results might only be realized in the longer term (akin to pulling up a new plant to see if the roots have taken). Building a learning system in government calls for more participatory approaches to evaluation or review throughout the policy process (in its design, data collection, and making sense of results).

Whether or not formal evaluations are undertaken, policy makers need to adopt an evaluation mind-set. That means always having a working hypothesis and theory of change and continually asking: how will we know if we’ve made a difference or not? What evidence do we need to gather? How will we monitor and measure what’s changed, for good and bad, and for who? They need to constantly test their assumptions as new evidence and insights are revealed throughout the process of policy design to delivery and back again – test, reflect, and iterate.

Embedding evaluation into decision-making processes, as recommended by the OECD, means adjusting upstream policy design. Good policy, and the advice that underpins it, requires a multiple lens approach to evidence: looking back, looking deep and looking forward (what I call hindsight, insight and foresight). Hindsight, or learning from past policy success and failure, is crucial. Yet, the old-style ‘policy cycle’ that still dominates most guidance and training for policy practitioners situates evaluation as a discrete stage - an add on or afterthought - at the end of the policy process. To challenge that, I’ve articulated an alternative framework in the 5D policy advice model (described in detail with animation here). I’m delighted the Irish government has adapted the 5D policy model into its policy guidance. In the 5D, I situate evaluation (and engagement) as continuous and connected activities throughout the policy advice development process.

The 5D policy advice model and the evaluation loop

Evaluation can, and should, happen at any or all stages of the policy process. Evaluation loops through each of the Ds in the 5D model, from understanding the challenge to testing the quality of advice, to delivering the final solution. Figure 1 indicates types of evaluation questions (indicative not exhaustive) as lines of inquiry throughout the 5D process. The 5-d Policy advice value chain Figure 1. The 5D policy model with evaluation lines of inquiry

Evaluation across the 5Ds

StageQuestions to ask
DemandWhat are we trying to address/achieve (current state, evaluation criteria)?
DiscoverWhat have we learnt from previous evaluations, here and in other organisations, sectors or jurisdictions?
DesignWhat's our theory of change? How can we continually test our assumptions? How might we monitor, test and iterate our options as we learn more or as the context changes?
DecideHave we delivered advice that helps the decision maker take a decision? How might we measure the quality of that advice, including recipient satisfaction?
DeliverHow will we know if, and when, we've made a positive difference? Was our process good? What did we learn from it? Have we achieved desired outcomes/impact?

How can we encourage, ensure and assure an evaluation mindset throughout the policy process?

Adopting an adaptable and scalable policy process model that includes evaluation as a continuous loop in the policy process, rather than as a step at the end, can help embed evaluation as an integral part of policy design.

Further incentives might include building evaluation into policy standards, such as Cabinet paper requirements (mandating evaluation criteria in proposals to Cabinet or other decision-making fora) and into wider policy quality expectations. Evaluation should apply to policy advice itself. Aotearoa’s policy quality framework provides criteria to test the quality of advice to decision makers. All government departments with policy advice appropriations must report policy quality scores annually to Parliament. Quality criteria can similarly serve as a framework to test and iterate policy proposals throughout the policy process (including via policy panels and peer review). As a light touch approach, the first iteration of the PQF included acid tests (riffing off the UK Policy Tests) to be used as a quick evaluation of the quality of policy advice immediately before advice is delivered to a decision maker. I’ve resurrected that idea in various masterclasses and shared below (Figure 2). It provides another reminder to ensure evaluation is included in advice.

Policy Quality Advice Figure 2. Policy advice quality tests

Embedding evaluation into policy processes also requires some supporting infrastructure. Avenues to consider include:

- Tapping into system centres of expertise: The OECD survey reviewing progress on the 2022 recommendation, showed an increase in evaluation centres of expertise in member countries. Policy professionals can tap into that expertise both as sources of evidence and evaluation advice. In the UK, the well-established ‘What Works’ centres provide repositories of evaluation evidence in key policy areas. They have developed into an international network and shared source of evidence. In Australia, the Australian Centre for Evaluation (ACE), is actively building evaluation activity and capability across the federal government.

- Build organisational capability: Within organisations, policy leaders can encourage evaluation by allocating resources to socializing and supporting the evaluation imperative and capability. Many larger organisations have evaluation specialists on staff, but they are often only involved at the end of the policy process to evaluate results (or commission evaluations) of decisions already taken. How might they be more effectively involved in the policy process upfront, helping to define evaluation criteria, robust theories of change, and good data collection. By walking alongside policy staff, they can help boost the evaluation skills of those staff and improve the quality of policy advice and underpinning evidence.

- Build evaluation into policy guidance, methods and toolkits: There is a wealth of evaluation methods and tools. The key challenge is for policy practitioners to know the purpose of each method, when to use it, how to use it, and the relative costs and benefits. In the UK the Magenta book has been the evaluation bible (albeit underpinned by the traditional policy cycle). In Australia, ACE’s Evaluation Toolkit includes templates, tools and resources on how to plan, conduct and communicate evaluations. These guides need to be cross-referenced in policy guidance and toolkits, so they influence upstream policy processes.

- Enhance evaluation skills and mindsets: Evaluation should also feature in articulations of policy skills and expectations. The UK Policy Profession Standards include a rubric on evaluation as part of its ‘delivery’ pillar, while the Aotearoa Policy Skills Framework includes ‘monitoring and evaluation’ as a practice to support the delivery of quality policy advice. In a range of jurisdictions ‘professions’, with dedicated whole of government leadership and supporting communities of practice, are developing. Deliberate connections between the policy and evaluation professions would help to improve both the demand and the supply of evaluation activity and impact. Policy practitioners need to be intelligent customers of evaluation expertise. Most importantly, they need to develop their evaluation muscles and mindsets: like kids in the car on road trip, they need to constantly ask “are we there yet?” and know how to identify milestones and speedbumps along the way.

A call to action – looping the loop

More evaluation is not a panacea. Finite resources, short time frames and sometimes the reluctance of decision-makers to invite scrutiny of pet projects can act as disincentives to evaluation activity. But improving evaluation capability and considerations as part of upstream policy advice processes can help to support better evidence, better policy, and a more intelligent learning state. A systemic approach to evaluation has implications for ‘how we do policy’ and for public service capability. We need to boldly loop the loop.