‘Over the last half-century, statistical analysis has been adopted by every single scientific discipline, from archaeology to history to physics, as the empirical tool of choice to advance its understanding of the world. There is no going back. In this context, those who deny statistical analysis as an indispensable means to uncover and investigate the nature of reality are denying the need to advance and perfect the scientific method itself.’
Bill Luker, Jr. Ph.D.
‘I don’t really know any educated people, who claim that “(Statistical) sampling does not work for Big Data.”‘
Gregory Piatetsky-Shapiro
Myth #10: (Statistical) sampling does not work for Big Data (Volume)
Myth #11: Current statistics techniques will not work for Big Data
Data is an asset containing information. Extracting information from numbers with uncertainty requires the heavy lifting power of statistics. Big Data (Volume) has the potential to store more information and require more statistics.
We wish to deconstruct two of the many new myths that are being advanced by the current promotional hype: 1) sampling (statistical) does not work for Big Data (Volume), and the even more ridiculous claim that 2) none of statistics works for Big Data. This is like the cigarette ads of old, which claimed that smoking was good for your health; protects your throat; is good for your digestion; et al. Instead, people, and even their pets, developed all manner of cancers, including throat cancer.
In the case of Big Data, the ‘promotional industrial complex’ wants to sell things, things like books, magazines, conferences, workshops, new degree programs, software, advertisement space, and newly anointed ‘experts.’ New things sell better.
Myth: (Statistical) Sampling Does Not Work for Big Data (Volume)
Here is a statistics denial quote in the context of (statistical) sampling from a LinkedIn discussion: But you can not apply the same statistical [sampling] techniques to a 200 million rows data set, than to a 10 million rows data set, because of the curse of big data. Yes, we can. We are familiar with sample designs [techniques], some of which accommodate infinite populations. Done.
Myth: Current Statistics Techniques Will Not Work for Big Data
We are not sure how anyone could state this with a straight face. The purpose of this promotional claim is to help sell pretend ‘Big Data methods’ by fabricating an unmet need. This myth blends the misunderstanding that there is something other than statistics for solving uncertainty problems, with the suggestion that Bigness removes uncertainty anyway. This concoction is then rationalized with a ‘not invented here’ mandate, alternatively stated as, ‘it does not count until I certify it.’
Once volume, velocity, variety, and/or veracity become part of the data analysis problem, more statistical thinking and more statistical techniques are needed. The advances in analyzing Big Data fall mostly along four, decades-old, fronts in applied statistics: most effective approach, improved automation, faster computation algorithms, and better data reduction.
As to data reduction, the first step in any analysis is to get the data into an analyzable form. This often requires using sampling to reduce the number of observations and multivariate statistics to reduce the number of variables. During the reduction stage, we want to retain as much clean information as possible. The need for data reduction mushrooms as Big Data issues become a larger part of the data analysis. This is not only because of the size of the data, but also because data can be uncertain and redundant. E.g., consider the Big Data equivalent of one million identical jigsaw puzzles, each with one billion pieces50% missing.
Without statistics, there will be no analysis of Big Data. Those having problems with that are probably having the same problems with any data.
Close:
Statistics is all we have for analyzing Big Data and it works. Sampling reduces large numbers of observations while retaining as much information as possible. This capability makes sampling more important for Big Data.
The need for statistics was never to compensate for a shortage of data. It has always been to address uncertainty, which is present for Big Data. We provide a clarifying problem-based view of statistics in the May/June 2015 issue of Analytics Magazine, https://goo.gl/Wod3gk.
We sure could use Deming, right now. Many of us, who consume or produce data analysis, hang out in the new LinkedIn group: About Data Analysis. Come see us.
The entire Statistical Denial series can be found on Datafloq.
- Statistical Denial 1, Blog 1: Essays On Statistics Denial
- Statistical Denial 2: Statistics Debacles & The Coming Flood Of Statistical Malfeasance
- Statistical Denial 3: Applied Statistics Is A Way Of Thinking, Not Just A Toolbox
- Statistical Denial 4: Five Forces Pushing Statistics Expertise Out of Data Analysis
- Statistical Denial 5, Myth 1: Traditional Techniques Straw Man
- Statistical Denial 6, Myth 2: Why Statisticians Not Only Practice Within Traditional Statistics
- Statistical Denial 7, Myth 3: Are Data Mining and Machine Learning Distinct from Statistics?
- Statistical Denial 8, Myth 4: Why Prediction Is / Is Not Part of Statistics
- Statistical Denial 9, Myth 5 and 6: Why Statistical Significance Does Work For Big Data
- Statistical Denial 10, Myths 7, 8 and 9: 3 Statistics Denial Myths on the Volume of Big Data
- Statistical Denial 11, Myths 10 and 11: Debunking the Myth that None of Statistics Works for Big Data
- Statistical Denial 12, Myth 12: Publications Straw Man
- Statistical Denial 13, Myths 13 and 14: Minimizing The Profession