Posted in

Debunking the Myth that None of Statistics Works for Big Data

‘Over the last half-century, statistical analysis has been adopted by every single scientific discipline, from archaeology to history to physics, as the empirical tool of choice to advance its understanding of the world. There is no going back. In this context, those who deny statistical analysis as an indispensable means to uncover and investigate the nature of reality are denying the need to advance and perfect the scientific method itself.’

Bill Luker, Jr. Ph.D.

‘I don’t really know any educated people, who claim that “(Statistical) sampling does not work for Big Data.”‘

Gregory Piatetsky-Shapiro

Myth #10: (Statistical) sampling does not work for Big Data (Volume)

Myth #11: Current statistics techniques will not work for Big Data

Data is an asset containing information. Extracting information from numbers with uncertainty requires the heavy lifting power of statistics.  Big Data (Volume) has the potential to store more information and require more statistics. 

We wish to deconstruct two of the many new myths that are being advanced by the current promotional hype: 1) sampling (statistical) does not work for Big Data (Volume), and the even more ridiculous claim that 2) none of statistics works for Big Data. This is like the cigarette ads of old, which claimed that smoking was good for your health; protects your throat; is good for your digestion; et al. Instead, people, and even their pets, developed all manner of cancers, including throat cancer. 

In the case of Big Data, the ‘promotional industrial complex’ wants to sell things, things like books, magazines, conferences, workshops, new degree programs, software, advertisement space, and newly anointed ‘experts.’  New things sell better. 

Myth: (Statistical) Sampling Does Not Work for Big Data (Volume)

Here is a statistics denial quote in the context of (statistical) sampling from a LinkedIn discussion: But you can not apply the same statistical [sampling] techniques to a 200 million rows data set, than to a 10 million rows data set, because of the curse of big data.  Yes, we can. We are familiar with sample designs [techniques], some of which accommodate infinite populations. Done. 

Myth: Current Statistics Techniques Will Not Work for Big Data

We are not sure how anyone could state this with a straight face. The purpose of this promotional claim is to help sell pretend ‘Big Data methods’ by fabricating an unmet need. This myth blends the misunderstanding that there is something other than statistics for solving uncertainty problems, with the suggestion that Bigness removes uncertainty anyway. This concoction is then rationalized with a ‘not invented here’ mandate, alternatively stated as, ‘it does not count until I certify it.’   

Once volume, velocity, variety, and/or veracity become part of the data analysis problem, more statistical thinking and more statistical techniques are needed. The advances in analyzing Big Data fall mostly along four, decades-old, fronts in applied statistics: most effective approach, improved automation, faster computation algorithms, and better data reduction. 

As to data reduction, the first step in any analysis is to get the data into an analyzable form. This often requires using sampling to reduce the number of observations and multivariate statistics to reduce the number of variables. During the reduction stage, we want to retain as much clean information as possible. The need for data reduction mushrooms as Big Data issues become a larger part of the data analysis. This is not only because of the size of the data, but also because data can be uncertain and redundant. E.g., consider the Big Data equivalent of one million identical jigsaw puzzles, each with one billion pieces50% missing. 

Without statistics, there will be no analysis of Big Data. Those having problems with that are probably having the same problems with any data. 

Close:

Statistics is all we have for analyzing Big Data and it works. Sampling reduces large numbers of observations while retaining as much information as possible. This capability makes sampling more important for Big Data. 

The need for statistics was never to compensate for a shortage of data. It has always been to address uncertainty, which is present for Big Data. We provide a clarifying problem-based view of statistics in the May/June 2015 issue of Analytics Magazine, https://goo.gl/Wod3gk. 

We sure could use Deming, right now.  Many of us, who consume or produce data analysis, hang out in the new LinkedIn group: About Data Analysis.  Come see us. 


The entire Statistical Denial series can be found on Datafloq.  

  •  

Randy Bartlett, Ph.D. CAP® PSTAT® is a statistician/statistical data scientist with 20+ years of practice experience analyzing and reviewing data analysis; and leading business analytics teams.  He is currently a Business Analytics Leader at Blue Sigma Analytics.  He provides services for everything from strategic consulting for business analytics to data analysis and data management.  His services are reflected by his book, workshops, and presentations.

He designed 'A Practitioner’s Guide to Business Analytics' (McGraw-Hill, 2013) (https://tinyurl.com/jx8rcru) to be the foremost reference on how corporations can better implement business analytics and in this era of Big Data and the Internet of Things.  He discusses strategic topics, including culture, organization, planning, and leadership for business analytics, in Chapters 1-6 and in Day I of his workshop.  For tactics, he discusses Statistical Qualifications, Diagnostics, and Review; and Data Collection, Software, and Management in Chapters 7-12 and during Day II of his workshop.  He previously contributed to the Encyclopedia for Research Design and writes blogs (including a series on Statistical Denial), case studies (AIG, AstraZeneca, big pharma, Google Flu Trends, et al.), and articles (two in Analytics Magazine). 

Specialties: Leading quants; addressing Big Data; making and supporting analytics-based decisions; performing statistical review; evaluating datasets and software needs;and re-organizing analytics teams.

Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.