The question of quality vs. quantity has become so ubiquitous that it has arguably become clich. However, this long-standing issue remains a serious topic, especially when dealing with Big Data.
Traditionally we are asked to choose between quality or quantity one or the other. You cant have your cake and eat it too; or rather, you cant have your 5 supermarket cakes and single high-end bakery cake and eat all of them. OK, you could. But youd definitely pay for it the next day.
However, with modern data, this predicament has been largely resolved. While the main issue in the quality vs. quantity debate has traditionally been concerning the cost of storing massive quantities of data, today data storage is cheaper than ever, and decreasing every day.
Just think about this in the early 90s, youd find data storage costing $3,000 per GB (and thats a lowball number). Ten years later in 2000, it was $20/GB. Today were at $0.03/GB. The rise of cloud storage and habitual decline of downloading (at least in the classical sense) is helping to dramatically decrease the cost of storage.
Today its cheaper than ever to collect data good thing too, because there is a ton of data out there to collect. Were dealing with ever-increasing consumer web activity and easy means of collecting all kinds of data about web visits, site activity, click-through rates, conversion paths, and more!
With such vast deposits of data, the main issue becomes: what do we do with all this data? The value of a Big Data platform is contingent on what processing that data leads to. What sense can be made of the data? You could have a gazillion petabytes of data, but if no understanding can be siphoned, its all useless.
Does your Big Data analysis allow your business to make more intelligent business decisions? Does it lead to new product innovations? Can it help you lower costs? Conduct better transactions? Improve your ROI? Better understand customer behavior? Quantity is meaningless if it doesnt drive actionable value.
Big data isnt so different from a Monet painting. Examine your data too granularly and youll only see a bunch of dots and blotches before your eyes. Step back a bit for the bigger picture, and youll be able to make sense of your data at the macro level and decipher meaning. Still, youll obtain the best value when viewing from just the right angle and the perfect perspective.
When you get that right angle, that super sweet spot, the benefits can be massive. Companies relying on delivery systems are now utilizing GPS truck data to track delivery routes, speed, performance, and scheduling. UPS used this kind of Big Data to optimize routes and save massive amounts of time and money. In 2011 they saved 8.4 million gallons of fuel by trimming 85 million miles off daily routes. Considering that cutting only one daily mile per driver saves them $30 million, the cumulative savings are more than a little impressive.
Amazon is the ultimate example of the magic that can happen when Big Data is utilized properly. Amazon uses massive amounts of collected consumer data to predict what products customers are most likely to purchase after viewing specific items.
Its easy to see why large data quantities are useful and valuable. But do we still need to worry about data quality? Huge data quantities make individual errors largely unimportant. So long as the majority of data is accurate, a few corrupted pieces of data will be as inconsequential as a few specks on a windshield.
However, in some cases data cleaning is an essential step in filtering out erroneous pieces of data. In order to determine just how important quality is for your Big Data, its important to consider the cost of corrupted data, especially when bad data can have tremendous impact on individual scenarios. For example, as Jeff Kelly notes, when you are using Big Data to determine medicine dosage for critically ill patients, you need to be relying on good data, not just mostly good data.
Ultimately, the importance of data quality will depend on:
A.How much effort will be required to correct erroneous data?
B. What is the data being used for, and what are the consequences of using bad data?
If the results of a few bad data pieces arent dire, its probably not worth the parsing. If youre making decisions based on large quantities of data, youre likely to be fine with a few minor blips. Honing in on narrow data segments is when youll want to be more cautious and consider paying more attention to quality.
What do you think is more important when it comes to Big Data: quality or quantity?