In 2013, Krishnamurthy Ramesh, head of the tiger monitoring program at the Panna Tiger Reserve in central India, received an email warning him that someone was trying to hack into his email account. Ramesh oversaw the electronic tagging of the endangered Bengal tigers on the reserve, and kept the GPS coordinates of tagged animals on his work email.

Although the hack was unsuccessful and the data was encrypted, it raised questions about what could happen if the animals' whereabouts got into the wrong hands. The same data that allows staff at the reserve to keep track of India's dwindling tiger population would have allowed the hacker to track the tiger to its exact location. Such incidents of "cyber-poaching" have become more common in the last few years, as scientists build up data sets of endangered animals, revealing their precise locations.

This is one example of what happens when big data goes wrong. There are many others. Rather than make the picture clearer, data can throw back a distorted picture of reality. Although policymakers set out with the best intentions, making policy off the back of bad data is like building on a fault line. AI and machine learning can exacerbate existing inequalities, while opening up data for anyone to access can lead to unintended and disastrous consequences.

Here at Apolitical, we’ve found the horror stories to help you plot a route through the pitfalls.

The pitfalls of open data

Open data is supposed to get the right information into the right hands. In 2017, India’s government opened data from across its ministries to public access, and sparked innovations immediately. It led to the development of a unified payment interface called Open Street Maps, a full-stack web application framework called Frappe, and an open learning system called DigiLocker.

But sometimes, the unintended consequences aren’t good ones. In 2015, two Spanish citizens were arrested in South Africa’s Knersvlakte Nature Reserve. A ranger discovered them with a pickup truck full of endangered plants, while a further search of their guesthouse revealed hundreds more.

The Spanish couple had been using online data from JSTOR Global Plants, iSpot, and social media to track down endangered plants to sell on their website. Geotagged data revealing the location of the plants had led them directly to where they were hidden, allowing them to collect rare varieties and sell them to collectors. In a similar case, poachers scoured social media posts of South Africa’s Kruger national park for geotagging that revealed the location of valuable animals.

Taking power from the poorest

When governments want to increase efficiency and transparency, they shred the paper records and turn to a centralised digital database. From 2009, India started to give every citizen over the age of 18 a unique digital identity. The “Aadhaar” system allows 1.13 billion Indians to access government services, pay taxes and open bank accounts. Previously, many of these citizens had no identification at all, and couldn’t access bank accounts or government benefits.

The same process can be used to take power from the poorest. From 2001, the Indian state of Karnataka began to computerise its land ownership records. The “Bhoomi system”, lauded by the World Bank and local politicians, made small landowners replace their existing land title with a standardised record, which was uploaded to a central system. Local village offices were closed and administration shifted up a step to district government, or “taluk”, level. Through this, the government of Karnataka claimed to increase efficiency and transparency, while enabling both swifter and clearer business transactions.

The result was the opposite. Far from reducing corruption, the Bhoomi system provided an opportunity for local government agents to get involved in bribery. Before, small landowners would pay small bribes to village-level officers to speed up the administration of a land title. Following the reforms, they had to spend several bribes to navigate the four administrative layers which replaced the local officer. While the local bribes could be negotiated, the new system allowed agents to systematise bribery, and collude amongst each other to set a market rate. These layers also slowed down the process for the small farmers significantly.

But the most damage was caused by the computerisation of the land records itself. By standardising land records and uploading them to a central platform, the Bhoomi system gave the people in control of the data a bird’s eye view of the land in Karnataka. This structured the land market, making it easier to plan large projects, and for developers to bid for them, but at the expense of the small landowners. Whereas before small and medium-sized farmers had access to their physical land records, computerisation took control from their hands and gave it to local government, multinational developers and big farmers. Computerisation was used to strong-arm land from their possession and reinforced existing poverty.

Discriminatory algorithms

Automation can save lives. By training algorithms to find patterns in masses of data, we can automate complex processes and predict everything from the likelihood of fires to child abuse from calls to child protection hotlines. Yet such algorithms are built on historical data, and if these contain biases, the algorithms they are based on can exacerbate them.

US courts in states such as Idaho and Colorado use an algorithm to measure each convicted criminal’s chances of reoffending before they are sentenced. Criminals are asked to fill out the LSI-R questionnaire, which asks them questions ranging from whether they had previously been involved in crime, to where they live, to whether anyone in their friends or family had been involved with the police. Based on correlations and previous case data, the model outputs a score, from high-risk to low-risk, which rates the criminal’s likelihood of reoffending.

These systems aren’t neutral. While they are designed to prevent the bias of the human judge, they can’t bypass social biases in the data. In the US, black people and ethnic minorities are more likely to have been stopped by police and to have family members who have committed crimes. Since the model works off previous involvement with police, rather than crime, simply being stopped and searched can impact a score. If treatment of minorities isn’t equitable, then the output score is distorted. These metrics often serve as proxies for race or poverty - they don’t judge the person, but their circumstances.

Technology brings new opportunities - but remember the human element

Opening data has led to positive innovation, but it can also lead to exploitation. Digitisation can provide the opportunity for the strong to take from the weak. Algorithms can do the work of whole departments in a fraction of the time, but they also exacerbate discrimination.

All of these dangers can be avoided, but it requires awareness of the flaws in data. Data isn’t neutral, and contains biases. Spot these, and your policies won’t exacerbate them. Working with data can even help to reveal them in the first place.

Good policymakers will be aware of these fault lines. When people get to grips with any new technology they discover the potential for exploitation alongside improvement. Remembering that human beings are always part of the equation is the first and most essential step when working with data in government, and understanding it can cause the biggest headaches.

(Picture credit: Flickr/beachmobjellies)


Make sure to share your own thoughts with the author by leaving a comment below