This article is written by Andrew Ford, Assistant Manager, International Strategy – Innovation and Science, Australian Government Department of Industry, Science, Energy and Resources.
The views expressed in this article are entirely those of the author and do not necessarily represent the official policy or position of the Australian Government or any of its agencies.
Innovation rankings are really popular, and not without reason: we like to know how well our department, industry or even country is doing, especially when we’re doing well. These rankings usually take the form of league tables based on an aggregation of multiple metrics into a single number. Many criticisms have been raised, but they retain popularity because they promise a simple and accurate ranking of their subjects, mainly countries or universities. Naturally, time-poor politicians and bureaucrats are attracted to this promise of finding what they need to know about a nation’s innovation or competitiveness, or university quality, in one simple ranking.
Unfortunately, this promise is not just a mirage, but a dangerous fallacy, which threatens to distort policies and programs to the detriment of taxpayers.
- Want to write for us? Take a look at Apolitical's guide for contributors
Composite rankings include the Global Innovation Index (GII), Global Competitiveness Index (GCI), and several World University Rankings. The GII and GCI take multiple indicators, and aggregate them into “pillars”, which are then further aggregated into the final ranking.
Despite their value, their flaws are many, and I will cover only the most significant. I will concentrate on the GII as I am most familiar with it, but similar points apply to all of them.
The biggest issue is that the whole concept of aggregating disparate indicators into a single final number is fundamentally unsound
It is important to note that the producers of the GII are not blind to its limitations. The introduction to the 2019 GII says, “The GII is not meant to be the ultimate and definitive ranking of economies with respect to innovation. Measuring innovation outputs and its impact remains difficult…” They have been addressing the flaws over time, although this means that comparisons from one edition to the next are not valid. And despite their efforts, major issues remain.
The greatest risk is that users, with no time or expertise to critique the methodology and conclusions, are using the GII as “the ultimate and definitive ranking of economies with respect to innovation”, despite the advice of its authors.
Invalid Concept
The biggest issue is that the whole concept of aggregating disparate indicators into a single final number is fundamentally unsound. Even if every component was accurate, robust and meaningful to innovation performance, there’s no testing of whether the combination of factors is better when the average rating on each is higher – it’s possible that a few are disproportionately important, some are irrelevant, or even that being “good” in several of the indicators simultaneously is actually “bad” for innovation performance overall.
The choice of indicators also biases the result towards a particular view of what makes a country innovative. This has a pernicious effect on innovation policy by encouraging governments to all strive towards the same model of innovation, when we don’t actually know it’s correct, or even that there is a single correct model.
Volatility
There is severe volatility in the rankings from year to year. For instance, Singapore came in 3rd in 2012, 8th in 2013, 5th in 2018, and 8th in 2019. The 2019 GII acknowledges that “scores and rankings from one year to the next are not directly comparable”, yet governments trumpet a rise of one place and are attacked for a fall of one.
Factors that contribute to volatility include:
- sensitivity of any ranking to meaningless small fluctuations in data,
- data coverage and availability,
- treatment of outliers and missing values,
- the sample of economies in a given year, and
- adjustments made to the GII framework methodology.
Two particularly bad problems are the reference year and normalisation. Quoting from the 2019 GII:
“The data underlying the GII do not refer to a single year but to several years depending on the latest available year for any given variable. In addition, the reference years for different variables are not the same for each economy. … The timeliest possible indicators are used for the GII: 37.3% of data obtained are from 2018, 33.3% are from 2017, 9.3% are from 2016, 4.8% from 2015, and the small remainder of 5.3% from earlier years.
Most GII variables are normalized using either GDP or population ... Yet, this implies that year-on-year changes in individual variables may be driven either by the variable’s numerator or by its denominator.”
I approve of normalisation, particularly by GDP and population. It is nonsense to compare China to the Seychelles otherwise. Yet the way this is done for the GII means that apparent performance changes can be driven by non-performance issues like population growth or decline. Since the reference year is not consistent, countries are also not being compared accurately to each other at the same point in time.
Data availability and source
Missing data is handled problematically. In the 2015 edition it is noted that “To be included in the GII, economies must have a minimum data coverage of … 60% … Missing values are indicated with ‘n/a’ and are not considered...” This can lead to ridiculous results. For instance, eSwatini beat Australia one year for “Knowledge and Technology Outputs” when half the Swazi data was n/a and just one of the other fields had it ahead of Australia, but so far ahead it outweighed all other data!
It was ahead in that one field because it’s a tax haven, not due to any real innovation edge. The 2019 edition has raised the bar to 66% coverage, which is still far too low – one third of the data for a country may be missing, which can seriously distort the comparisons. More than 13% of countries have coverage lower than 75% in the “Output” sub-index.
The GII correctly notes that “innovation surveys… fail to provide a good and reliable sense of cross-economy innovation output”, but still uses surveys! For 2019, “a total of 57 variables are hard data”, an improvement over earlier editions. However, “a total of 18 variables are composite indicators and 5 are survey questions from the World Economic Forum’s Executive Opinion Survey”, still too many questions rely on too narrow an evidence base — for Australia “survey” means asking members of a single business lobby group for their opinion.
Quirky Indicators
The GII takes the average score from the most controversial of all the university league tables for just the top three universities in a country, no matter how many universities that country has – clearly an inappropriate way to compare Singapore to China. Some form of normalisation would be better.
For relative investment in R&D (Gross Domestic Expenditure on R&D - GERD - as a percentage of GDP), higher is regarded as always better. Actually, there is strong evidence that too high can be harmful, due to diminishing returns in R&D and opportunity costs in other areas.
I encourage everyone to approach these composite rankings like Wikipedia – not as the repository of all truth, but as a collection of useful links to sources of real knowledge
GERD Financed by Abroad may actually reflect the foreign investment rules in a country, not its innovative capacity.
The H-Index, another poorly constructed “fit everything into one number” indicator, is used as an indicator of knowledge creation.
My favourite quirky indicator, however, is the cost of redundancy dismissal, defined as:
Sum of notice period and severance pay for redundancy dismissal (in salary weeks…)… The worker is a cashier in a super-market or grocery store.
Labour mobility is important to innovation, but is this an appropriate indicator? A lower number ranks better in the GII, so it’s regarded as innovation-friendly to be able to dump staff cheaply, which is questionable. It tells us nothing about the more important side of mobility – ease of gaining a new job. There is nothing covering that, even tangentially, in the GII. The idea that the notice period and severance pay for a cashier indicates the appropriateness of the regulatory environment for encouraging innovation is deeply unconvincing. This was clearly a case of taking whatever data was available, but it probably wasn’t intended for this particular use by the original collectors.
How to use rankings
I encourage everyone to approach these composite rankings like Wikipedia – not as the repository of all truth, but as a collection of useful links to sources of real knowledge, served up with a few bits of nonsense. — Andrew Ford
(Picture credit: Unsplash)

Log in or sign up to continue the conversation