Back in the old days of Web 1.0, there has been a short-lived debate on the trustworthiness of information found on websites. We were taught to check that certain information was referenced by other websites as well, that the sources were present and legitimate, and that comparable information could be found on other independent websites. Despite these precautions, a Dutch journalist argued that it would easily be possible to set up a group of websites with misleading information, that would all reference each other and would subtly draw a complete web of misinformation. It would be a matter of time before the information would be picked up by mainstream media and would consequently turn into truth. (Article in Dutch, and the article translated by Google.)
Then the Google PageRank entered the scene, and the idea behind this algorithm was so powerful that people quickly assumed that when a website made it to the top 10 of any Google search, it would be trustworthy. Although misleading information still exists, powerful interplay between Google and Wikipedia nowadays leads most internet users towards information that has a minimal degree of trustworthiness.
How is that for data? How do we know that the profound Big Data-based insights that we base on wildly separate and variate datasets can be really trusted? Could we think of a PageRank for data?
These days, as a kind of rebound to Big Data evangelism, we see a growing list of examples where Big Data had it all wrong. Just as humans, it turns out that data can be biased, as well. Therefore, it would be good to have data due diligence procedures, like proposed by Kate Crawford.
But its not just the data itself. The way these data are processed could lead to wrong or biased results as well. Thats why in this article the algorithmic accounting reporting approach is proposed; thats why Cukier and Mayer-Schnberger propose a new profession of algoritmists that are able to dissect, understand, explain and validate data processing algorithms.
Transparency is a great concept. The power behind open source and wikis is Linus Law, that states given enough eyeballs, all bugs are shallow. This law often leads to the idea that a complete transparent information exchange will inevitably make it impossible for misconceptions to stay in the mainstream discussion. This line of thought was already expressed during the first days of the telegraph, where it was predicted that journals would not have the possibility anymore to create a hype, because the Electric Telegraph would very quickly deliver the real facts, as James Gleick memorizes in his book The Information’. The telegraph, the telephone, the television, and internet have all been expected to be able to end wars. On a slightly different track, recently, Bono suggested that Big Data would end poverty in the world. Although I like being an optimist, history and reality somehow tell me that just transparent exchange of information alone is not the whole solution. Transparency (of data, of algorithms) is not enough.
When a system is too complex to understand, transparency will not help us not even with the most skilled algoritmist to explain what is going on. We should design our data chains, our systems and our algorithms in such a way that we can trace what has been going on. Besides a critical review of our data sources, and besides transparency of the algorithms that feed on it, we should also be able to design for accountability.
The goal to design for accountability contains three challenges.
First, it is a technological challenge. Every computer programmer that has been chasing an elusive software bug for days and days, knows that it can be very difficult to track down a specific system feature that explains its behavior. For larger systems, there is a combination of approaches that include log files, modular design, fault tolerant behavior and other patterns. For data-intensive systems and algorithms we need to get to comparable patterns.
Then, there is the policy challenge. Even apart from legal issues, it is already difficult enough to organize the web of actors in any reasonably complex system. When organization (or department, or module) C receives wrong information from B, will B then be held liable? Or could it be traced further back to A? Even if we are able to trace down data sources, we will very often find that the data is provided as is (as could be the case with Open Data), or we cannot verify the status of the data source itself (is the sensor calibrated; was the transaction legitimate).
Lastly, the psychology around information systems is not developed enough for data-intensive algorithmic systems. Reactions of the public towards complex ICT systems are very unpredictable. Two examples serve to illustrate this. When a number of Dutch banks became the target of DDoS attacks (mid 2013), many people started to doubt whether their money was safe or not. Another example is the reaction towards two different privacy-related services in two countries: in the Netherlands, there was quite some objection towards the smart energy meter because it was perceived as a privacy-threatening device, where in Germany this debate was much less intense. On the other hand, in Germany many people objected the Google Streetview service (demanding that their houses should be pixeled out because of privacy concerns), where in the Netherlands this service was embraced without much discussion. The conclusion is that we just do not (yet) have the frame (knowledge and vocabulary) to effectively deal with data-intensive algorithmic systems. Systems that are designed for accountability could greatly help the public to develop such a frame.