Posted in

How Can Blind Spots in AI Help Foster Online Privacy?

Machine learning application has now made it possible to detect cancer cells and create collision-proof self-driving cars. But, at the same time, it also threatens to turn over our notions of what’s hidden and visible. 

For instance, it enables the highly accurate facial recognition, sees through the pixelation in photos, and even uses data available on social media to predict sensitive traits like an individual’s political orientation, as was the case seen in the notorious Cambridge Analytica scandal. 

These same machine learning applications suffer from a peculiar sort of blind spot, which usually humans don’t do. This blind spot is a fixed bug which can make an image classifier mistake a rifle for a jet plane, or create an autonomous and free vehicle by a stop sign. All these misclassifications are known as adversarial examples, have been seen as an irking and severe weakness in several machine learning applications. Only a few small tweaks to an image or some additions of decoy data to a database can easily fool a system to end up entirely wrong conclusions. 

Researchers have suggested that attackers are increasingly using machine learning to compromise on user’s privacy, as demonstrated by the increasing complexity of cybercrimes, such as phishing and malware attacks.

However, now, this vulnerability can be used as a weapon to defend other people’s privacy. Efforts have been made to structure a method for exploiting the adversarial examples to make the re-identification and de-anonymization attacks less efficient and operational. 

Let’s find out more about the efforts by which blind spots can help in increasing online privacy.

A Sprint of Fake Likes

The data science firm paid a few dollars to thousands of Facebook users for providing answers to various personal and political questions. The given answers were linked with their public Facebook data. All this was done to create a set of training data. Later, the firm trained a machine learning engine with that data set. The resulting model can predict private political opinions and beliefs based on public Facebook data. 

To further test the hypothesis, a group of researchers used a corresponding data set, i.e., reviews within the Google Play store. They collected hundreds and even thousands of ratings in the Google app store. All these ratings were put forward by the users who have also shared their location on a Google Plus profile. 

A machine learning engine was trained with the data set to predict the home town of the users based on their ratings. The research concluded that based on Google Play likes, some machine learning techniques could recognize the city of users with around 44% of accuracy on the first attempt. 

The researchers tried to break this machine learning engine with adversarial examples. After modifying the data in different ways, it was found that adding a fake app rating that was chosen to spot an incorrect city statistically produces a small amount of noise. The noise reduces the accuracy of the engine’s prediction back to the random guess. Hence, within few changes, an attacker’s efficiency is reduced, and a user’s profile can be protected. 

However, the game of predicting and protecting user data doesn’t end here. If the attacker is aware of the fact that adversarial examples might be protecting data from the analysis, then they can generate their adversarial examples to include in a training data set. 

By doing so, the machine learning engine will be harder to fool. The privacy defender can respond it by adding some more adversarial examples to stop any more robust machine-learning engine which results in an endless counterattack. Even if the attacker uses any machine-learning, by adjusting the adversarial examples all such methods can be prevented which aimed at breaching user’s privacy. 

To Monitor a Mockingbird

Another research group also experimented with a form of adversarial example data protection which intended to resolve the prediction and protection game. The researchers looked at how adversarial examples can prevent a possible privacy leak in tools like VPN and the anonymous software Tor, mainly designed to hide the destination and source of the online traffic. 

Privacy invaders who gain access to encrypted web browsing data in transit can use machine learning to identify patterns in the traffic, which allows a spy to predict which website and even specific pages an individual is visiting. Researchers in their tests found that the technique famously known as web fingerprinting can recognize a website among a pool of 95 possibilities with almost up to 98% accuracy. 

The researchers identified that they could add adversarial examples like noise to encrypted web traffic to avoid web fingerprinting. However, they went further and attempted to short-circuit an adversary avoidance of those protections with adversarial training. 

To accomplish this, they generate a complete mix of adversarial examples tweaks for a Tor web session. In this session, a collection of changes to the traffic designed not to trick the fingerprinting engine into a falsely detecting website’s traffic as another, but rather than combing adversarial example changes from an extensive collection of decoy’s website’s traffic was done. 

The resulting system, which the researchers called a ‘Mockingbird” because of its combined mimicking strategy, adds a significant overhead of approximately 56% more bandwidth than the regular Tor traffic. It makes fingerprinting even more difficult. The accuracy of the machine learning model predictions was dropped to between 27% and 57%. Due to the randomized way of tweaking data, it becomes extremely tough for an attacker to overcome adversarial training. An attacker won’t find it easy to come up with all the possibilities to invade the user’s privacy. 

Wrapping Up

The above-explained experiments use the adversarial examples as a protective mechanism; instead, as a flaw is much promising from a privacy point of view. Also, the majority of academics are working on machine learning views adversarial examples as a problem to solve instead being a mechanism to exploit.

Enthusiastic Cybersecurity Journalist, A creative team leader, editor of PrivacyCrypts.

Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.