Posted in

Why Statistical Significance Does Work For Big Data

“The key element for a successful (big) data analytics and data science future is statistical rigor and statistical thinking of humans”

Diego Kuonen

#5: For a large number of observations (Big Data: Volume), all the variables are significant so statistics does not work

#6: Statistics does not accommodate ‘consequential’ statistical significance

Myth #5 builds upon the old confusion around significance testing that comprises this second ‘ancient’ myth (#6).  Suppose that you are building a predictive model based upon a billion observations (n) and using 500 variables (p).  After you press the magic ‘make model‘ button, all of the parameters are ‘significantly different’ … from zero.  Conclusion, statistics does not work … WRONG. 

Here is a quote typifying the misunderstanding:

‘One big reason [why statistics does not work] is that everything passes statistical tests with significance,” he says. If you have a million records, everything looks like it’s good [significant].” According to the same person, there’s a difference between statistical significance and what he calls operational [consequential] significance.  

First, if you are building predictive models, then why are you using hypothesis testing?  Second, why are you using hypothesis testing?  Third, if you must use hypothesis testing, then try the right one.  Use the ‘new’ breakthrough from Neyman-Pearson (ca 1932), which addresses (consequential) statistical significance.  Now we will be more specific about these three disconnects. 

 

First, if you are building predictive models, then why are you using hypothesis testing? 

For predictive models, whether coefficients are significantly different from zero is not the primary consideration.  The point is whether the model predicts.  We know, you seek parsimony by dropping parameters, which have coefficients that are not significant from zero.  Hint: There are better statistical avenues to parsimony; use statistics designed for that task.   

Recall that there are four modeling objectives: coefficient estimation, prediction, grouping, and ranking.  Hypothesis testing was conceived for decision making largely in the context of coefficient estimation.  As such, it is only an important sideshow to the main show of statisticsthe logic of numbers with uncertainty. 

 

Second, why are you using hypothesis testing? 
Confidence intervals generally have more utility than hypothesis tests.  We know, sometimes you just want or need a hypothesis test, yet not for prediction.  Also, confidence regions nicely address multidimensional needs. 

 

Third, if you must use hypothesis testing, then try the right one.  Use the new breakthrough from Neyman-Pearson (ca 1932), which addresses (consequential) statistical significance. 

Now we have arrived at Statistics Denial Myth #6, the old confusion between the Fisherian school of hypothesis testing and Neyman-Pearson.  Fisher was the first to address the matter of hypothesis testing and he developed a logical approach, which compares unknown parameters to zero.  Neyman-Pearson hypothesis testing famously expands this work by insisting on an alternative hypothesis and adding a term,, as a cutoff for (consequential) statistical significance.  This has been called practical significance, economic significance, etc. and now operational significance.  This portmanteau hypothesis test allows coefficients to be compared to any value,

For example, suppose that if a coefficient exceeds some consequential value , then retaining it is statistically significant. The hypotheses might take the following form:

,

where  is the cutoff for a consequential difference.  Neyman would say that the alternative hypothesis, , should represent the consequential scenario.  (See ‘Encyclopedia of Research Design,’ Vol. I, SAGE (2010), p. 298).  Hence, consequential significance is statistical significance.  As always, see a professional for your advanced statistics needs. 

Close:

Confusion about hypothesis testing is completely understandable, yet not acceptable for self-professed experts. 

While these misunderstandings have an amusing side, they also have an edge.  At the extreme, we have seen hucksters broadcasting mischaracterizations of statistics to better position their lesser qualifications or to blunt legitimate criticism of their blatant mistakes.  One common claim coming from hucksters innocent of statistics is that we do not need statistics anymore because we have access to them. 

In Blogs 2 & 3, we discussed the harm caused by promotional hype extreme enough to adulterate statistics and circumvent best practiceour best tools for extracting the information. 

We sure could use Deming, right now.  Many of us who embrace the explicit rigorous logic and protocols of these tenets of data analysis hang out in the new LinkedIn group, About Data Analysis.  Come see us. 


The entire Statistical Denial series can be found on Datafloq.  

  •  

Randy Bartlett, Ph.D. CAP® PSTAT® is a statistician/statistical data scientist with 20+ years of practice experience analyzing and reviewing data analysis; and leading business analytics teams.  He is currently a Business Analytics Leader at Blue Sigma Analytics.  He provides services for everything from strategic consulting for business analytics to data analysis and data management.  His services are reflected by his book, workshops, and presentations.

He designed 'A Practitioner’s Guide to Business Analytics' (McGraw-Hill, 2013) (https://tinyurl.com/jx8rcru) to be the foremost reference on how corporations can better implement business analytics and in this era of Big Data and the Internet of Things.  He discusses strategic topics, including culture, organization, planning, and leadership for business analytics, in Chapters 1-6 and in Day I of his workshop.  For tactics, he discusses Statistical Qualifications, Diagnostics, and Review; and Data Collection, Software, and Management in Chapters 7-12 and during Day II of his workshop.  He previously contributed to the Encyclopedia for Research Design and writes blogs (including a series on Statistical Denial), case studies (AIG, AstraZeneca, big pharma, Google Flu Trends, et al.), and articles (two in Analytics Magazine). 

Specialties: Leading quants; addressing Big Data; making and supporting analytics-based decisions; performing statistical review; evaluating datasets and software needs;and re-organizing analytics teams.

Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.