Earlier this week in Washington, D.C., the Future of Privacy Forum and Stanford Law Schools The Center for Internet and Society presented a one day conference “Big Data Privacy: Making Ends Meet.” That conference showcased numerous papers by luminaries in the privacy community. There were several common themes among the papers. One of these was the need to place contractual limitation on the transfer of de-identified Big Data sets. The argument being that public release of data, without limitation, allows “bad” actors (be they criminals or academia) more leeway and access to re-identify participants. This is one of the suggestions of Ira Rubenstien for his proposed “A Pretty Good Privacy Solution.”
However, unlike Phil Zimmerman’s Pretty Good Privacy software whose name was chosen tongue in cheek because it really did provide the best email encryption option of the day, Rubenstein’s offer is only marginally good. Particularly important is that one of the benefits of Big Data is the ability to find patterns without knowing what you’re looking for. Public release of data allows for a wider selection of individuals and groups to cull the data for interesting information. Restricting the data to vetted parties only serves to limit the analysis. Perhaps, more importantly, from a privacy perspective, the access control approach is just not that novel. Certainly, restricting access provide additional privacy to the data subjects but it is a known solution.
One of the more intriguing paper presented was by Arvind Narayanan and Jonathan Mayer on Privacy Substitutes. Traditionally, privacy competes with commercial value in making decisions on the collection and use of information. However, there exists technology that can provide positive sum approaches (principle 4 of the Privacy by Design principle). These privacy substitutes increase both privacy preservation and commercial value of the system. One example, relevant to Big Data, is the use of differential privacy which introduces quantifiable noise into the data set. This prevents privacy invasive queries directed at specific individuals or groups but still allows broad queries to tease out patterns in the data. Unfortunately, many organizations fail to implement such technologies out of either ignorance or the difficulty in balancing the notion of adding errors with system design principles.
A paper by Ryan Calo notes the possibility of Consumer Subject Review Boards for privacy similar to the ethical board that govern biomedical and behavioral science research. This is a concept previously brought to my attention by my frequent co-conspirator Dr. Stuart Shapiro. Calo emphasizes that while a direct correlation can not be made there is much similarity between behavioral research in academia and behavioral research in commercial entities trying to identify new ways to convince people to buy their products. The ethical considerations need to be similar as well, with a focus on minimizing the harm to the subject and maximizing benefit. Big Data firms might do well to create their own oversight boards similar to the Institutional Review Boards that govern research on human subjects. Calo call these Consumer Subject Review Boards.
I, too, submitted a paper to the conference but alas it was not chosen. However, the reader of BigData-Startups needn’t fear. You got a sneak preview of my paper in my blog post several months ago. Despite my paper not making the cut, there were many excellent articles that anybody interested in Big Data and Privacy should check out.