Posted in

Will Sensor Data Drive The Big Data Development?

Its hard to start any presentation on Big Data without some mentioning of at least the Volume V out of the 3Vs (or how many Vs you have thought up yourself additionally). Or to treat Big Data without indulging in the Petabyte Exabyte Zettabyte – Yottabyte mantra . But where do all those bytes come from?

Of course a company embarking on the Big Data train will start with what it has. Your own CRM and transaction data, maybe a repository of documents here and there, and of course social media. Most of this data is human-generated. This is not surprising, given that for most office jobs today, a large part of your working life is devoted to looking at screens and typing more data into some system.

However, machine-generated data will quickly dwarf human-generated data. Thanks to Moores law that still seems to be valid, and thanks to IPv6 that enables us to address every grain of sand in the world , we can expect that a large part of the phyiscal world around us will eventually somehow be online. Enter the Internet of Things. Sensors become a large part of these things.

What will happen when sensor data will be even more ubiquitous than human-generated data? What will happen when this sensor data will be as freely available (free as in beer) as human generated data? Could we compare these questions to the questions that we asked ourselves in 1993, when Netscape opened web 1.0 for us? Google was many, many years in the future then.

It is interesting to see how current activities might foreshadow the things that are to come.

Will Xively (formerly known as Cosm, and Pachube before that) be the next Google? They certainly try to gain mass by combining as much as possible sensors in their systems. However, we learned from Google that just the basic functionality alone does not make you a winner: AltaVista was first, but did not win. Google won because of quality of search results, but maybe even more because they set up an ecosystem.

We will also see an ecosystem around sensordata. Its not just the procurement of the sensor itself, or the installation, or the connectivity: thats usually done by the individual or organisation that somehow likes his or her sensor to be online. Its also not just about the API that enables the sensor to go online or the searching and indexing of sensor data. It will also be about the party that is able to do sensible things with all these sensor data: image interpretation, sensor data aggregation, visualisation.

And that will not be all. Although we do not yet see very clearly how things will develop, it can be expected that we will need some kind of trust provisioning. On the internet, nobody knows youre a dog, but how does this hold in the Internet of Things? It is instructive to see what happened during the Fukushima disaster: in the Pachube, as it was then still called, sensor ecosystem there turned out to be online Geiger counters in Japan at the time. They attracted a lot of attention. Many people tried to make sense of the readings, until it was remarked that some devices may be stored in university basements and consequently would show much lower readings that those devices placed outdoors. (Eventually, a huge crowd sourced movement began bringing Geiger counters online together with information on their placement.) How do we know that a specific sensor reading is trustworthy? Will we see some Google page rank lookalike?

Another issue might be that people take control of their data. How do we define data ownership? Will it be possible to, say, share your location data for free with an organisation that uses it for some societally beneficial goal, but to charge money for a company that just wants to make a profit out of your data?

We might need a model where data will always be accompanied by metadata. This would enable us to trace data back to its origin, thereby giving some clues on the trustworthiness of the data. It would also enable data providers to precisely track and control how their data is re-used. And finally it would very nicely match with the sticky privacy policy approaches.

It could very well be that the development of the way we use sensor data will lead the way we need to deal with data in general.

Image: Wikimedia Commons

Freek has a background in Digital Pattern Recognition (many years ago). He has been working on applications in Optical Character Recognition in the 90's (for KPN Research) and after that he has become involved in various ICT research projects. Freek has become increasingly interested in the relation between innovative technology and human/organizational behavior. In 2003, he joined TNO where he is now for some time the coordinator of the Big Data knowledge development program within TNO.

Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.