This article is written by Lucas Ribeiro Prado, is a public servant, writer, data scientist, ILO Evaluation and Monitoring specialist, Civic Innovation Ambassador of OKBR and co-founder of Meryt.


Ever since Harvard Business Review hailed the job of data scientist “the sexiest profession of the 21st Century”, it has been well and truly hyped.

But 8 years and many covers later, the real world, permeated by social and cultural inequalities, shows we need to take two steps back. This means talking about data literacy and the challenges it poses the public sector.

• Want to write for us? Take a look at Apolitical’s guide for contributors

MIT's definition of data literacy is the "ability to read, work, analyze, and argue" with information. We are not talking here about linear regression, decision trees, Bayesian inference, Convolutional Neural Networks or Adversarial Generative Networks.

Rather, we’re talking about basic notions of mathematics and statistics, such as outliers, standard deviation, median and causality, combined with computer knowledge.

In short, data literacy begins by being able to perceive the value data can have in a given context, and know how to use it to solve real problems. But to do this, one must know the right questions to ask and how to formulate them in order to obtain assertive and reliable answers.

Although the word “statistics” has emerged from the State itself for demographic purposes, today there are so few data scientists working in the public sector

The use of data to provide insights is nothing new. In 1854, in the face of the cholera outbreak in the City of London, physician John Snow managed to save thousands of lives just by mapping the occurrence of cases and crossing the data to discover the origin of contamination.

There were no AI, no cloud computing, no supercomputers and no beautifully graphics in videowalls. Today, we see world leaders making efforts to find as much data as possible about Covid-19 to help them make decisions.

But data does not speak for itself. To generate value, it needs to be interpreted, understood, translated, argued, communicated, shared and used. The excessive volume of data can end up being a problem, especially if there are no professionals capable of interacting with big data.

Unfortunately in the public sector, data scientists represent only 4.2% of the total number of professionals with a profile at Kaggle, the largest community of data scientists on the web. It is curious to note that, although the word “statistics” has emerged from the state itself for demographic purposes, today there are so few data scientists working in the public sector.

Not by chance, the report The Human Impact of Data Literacy showed that almost a quarter (23%) of public servants in the countries researched feel so overwhelmed when confronted with the data that they end up avoiding completing their tasks.

Meanwhile, 45% confess to feeling oppressed and unhappy at work at least once a week when reading, working and analysing data, although they consider themselves apt for this function.

Unlike a big tech companies with the resources to invest more and more in teams of data scientists, the public sector has been left behind, with budget restrictions that bar it from hiring skilled professionals to the civil service. According to the survey conducted by ODI, the most requested skills by public employees in the area of data science are: data visualisation (68%), innovating with data (68%), measuring success (67%), data analysis (64%), making data more usable (64%). Based on this, they developed an interesting data skills framework.

In the US Government repositories alone, around 300 terabytes per day of energy source data, 800 petabytes per day of climate data, and over 225 exabytes per month of internet traffic are being stored. Spreadsheet software, which is still used by 31% of UK public sector organisations, (according to Public Sector Data Report 2019), is not enough to handle this massive volume of government data.

A UK Department of Transport (DfT) miscalculation in a 2012 bid cost the public over £40 million in compensation in the notorious West Coast Main Line scandal and ultimately resulted in the removal of the government officials responsible for the trivial spreadsheet error.

Most public sector entities are still in the early stages of their exploration of data analysis, in search of data-driven solutions and public policies

Now let's think of those public servants, especially in the poorest countries, who do not know how to use even a simple spreadsheet editor to elaborate graphs, extract insight for their daily work, or even perceive a hint of fraud, failure or manipulation in data processing. How many millions of dollars could be recovered by the public coffers avoiding small data errors?

When the teams of data experts are in tune with all the other public employees who act in direct contact with citizens, it is possible to do incredible things. An example is Newcastle City Hall’s analysis of NEET individuals that helps the local authority identify the children most at risk of not being in employment or education.

Another is the case of the Amsterdam fire department, which collected data from different sources (information on roads, railroads, buildings, neighborhoods, etc) and compared it with historical records to predict where, when, and how often fires occur.

Because this is a new concept, most public sector entities are still in the early stages of their exploration of data analysis, in search of data-driven solutions and public policies. The Australian Government developed a data literacy strategy in 2016. In the UK, the Government Digital Service has structured a data literacy program for public officials. Countries such as Estonia, Brazil and Canada have also been trying to move forward in this direction.

Simply increasing the number of data specialists in government does not guarantee that managers will be able to seize the opportunities of the data revolution. Nor can we expect all public officials to become unicorns capable of developing sophisticated data mining algorithms in the government's gigantic datalakes. It's not just about technology — it's about creating a data culture in government organisations that is capable of generating public value and social impact.

The good news is that the language of data has a low learning curve, which means it can be easily learned and practiced in many open source tools. In addition, there is a large community of data lovers willing to share knowledge voluntarily on the Internet. We still have a long way to go to make data more accessible and readable to all public servants, but we need to start with the ABCs.

You can test your level of data fluency through this link.

We want to hear what you think about this article. Pitch an article to us or send your comments to hello@apolitical.co

(Photo credit: UnSplash)