Posted in

Blog 12: Statistics Denial Myth #12, Publications Straw Man

Von Moltke the Elder said, “No plan of battle survives the first contact with the enemy.” Similarly, no plan of analysis survives the first contact with the data. Doug Samuelson

In theory there is no difference between theory and practice. In practice there is. Yogi Berra

Myth #12: Statistics is defined by the recent publishing activities of statistics professors

Implication: There are these nonstatistical techniques for data analysis, which are needed to cover gaps in statistics

The myth is that the limits of statistics are defined by its academic publications and that there is a massive gap in applied statistical capabilities, which conveniently needs to be addressed by those spreading this myth.  This mischaracterization is typified by blogs such as, “Data science [analysis] without statistics is possible, even desirable,” which could only be made true if data analysis was excluded from data science and they are not going to do that.  Such claimants go on to be utterly incapable of listing anything other than statistics for analyzing data.  Their claims are like a dentist writing that we should stop using mathematics because now we can use ‘multiplication.’ 

The harm is that this statistics denial leads to another round of adultering statistics by excluding Best Statistical Practicethe proper qualifications, diagnostics, review, thinking, assumptions, and tools.  If statistics is not fully leveraged in analyzing Big Data; building AI; or orchestrating IoT, then these endeavors will not go far.  Just as Six Sigma could only must part of statistics and was correspondingly limited.  It would be better to leave applied statistics to the applied statisticians. 

Let us make four points that tear into the misinformation and confusion surrounding this claim. 

1.  Applied statistics is defined by its problems and not by a toolbox corresponding to academic publications and classwork. 

We provide a clarifying problem-based view of statistics in the May/June 2015 issue of Analytics Magazine, https://goo.gl/Wod3gk. 

All numeric problems involving uncertainty with the numbers are statistics problems, which have the domain of application as a second dimension.  Figure 1 reminds us that statistics problems are two dimensional: the domain and the statistics. 

Applied StatisticsFigure 1: Applied Statistics Matrix From ‘A Practitioner’s Guide To Business Analytics.

Statistics problems require statistical assumptions to infer, rather than deduce, from the partial information and within the domain.  E.g., if we ignore missing data that is germane to solving a problem, then we are making implicit statistical assumptions about the missings.  For uncertainty with the numbers, there is no unique solution to be ‘deduced.’  Whatever we do, or do not do, about the partial information makes statistical assumptions and affects how we infer from the available information and the context provide by our domain knowledge.  We spent the last decades developing Best Statistical Practice, which addresses how to think about and solve these problems and how statistical tools perform in the field. 

2.  Solutions for applied statistics problems are not confined to published ideas or techniques. 

A great deal of applied statistics is never published.  This can be because the ideas are proprietary (it can be foolish to share treasured techniques); not amenable to the publication process; etc.  Even when published, it can be very difficult to find appropriate references, e.g., the subfield of sampling is notorious.  Furthermore, anti-statisticians, wishing to covet by limiting a field that they do not comprehend, overlook the older publications, which layout the fundamentals for thinking about data analysis. 

3.  Applied statistics is not excluded from using ideas published in other academic fields. 

New (or old) statistics ideas can come from anywherejust like any other applied fields.  Furthermore, statistics is a hub science.  It learns from other fields and contributes to them.  To quote a Bando (martial art) instructor, ‘If we see something another school like Taekwondo use[s] that we like, we borrow the technique and it becomes part of our discipline.’ 

4.  Rethinking/republishing statistics ideas in another field does not move the topic to that field. 

It is not uncommon for old statistics ideas to be rehashed in another field.  Usually these ideas are combined with emerging breakthroughs occuring at the time.  This should not be interpreted as a paradigm shift in how to analyze data, which renders all that we have previously learned to be null and void.  That is, publishing a problem or technique outside statistics does not remove the underlying statistics (assumptions), thinking, et al. of the problem. 

E.g., machine learning is heavily published inside and outside of academic statistics.  We can think of this effort as building a class of tools to be applied in different fields, such as data management and data analysis.  This is like building sophisticated pressure gauges that can be used for steam engines or for brain surgery.  Using the same class of tool does not mean that expertise in data management equates to expertise in data analysis.  You still need a foundation in applied statistics to perform data analysis, which can be seen clearly from the statistical reviewespecially when a needed review is not ever performed. 

Close

Myth #12 is a straw man mischaracterization of statistics and by implication, applied statistics.  As with previous reinvention cycles, there will be noise and talking heads will rename whatever they can. 

The harm is that this mischaracterization arbitrarily misclassifies statistics problems as deductive ones.  Thereby, bypassing Best Statistical Practice in favor of counterfeit nonstatistics and opening the flood gate to statistical malfeasance.  

We define applied statistics by its problemsnumeric problems involving uncertainty with the numbers within the context of a domain.  Applied statistics is all we have for addressing these problems.  We can not solve them with nonstatistics and we should not split applied statistics into disconnected corrupt copies of itself.   

We sure could use Deming, right now.  Many of us who embrace the explicit rigorous logic and protocols of these tenets of data analysis hang out in the new LinkedIn group, About Data Analysis.  Come see us. 

The entire Statistical Denial series can be found on Datafloq.  

Randy Bartlett, Ph.D. CAP® PSTAT® is a statistician/statistical data scientist with 20+ years of practice experience analyzing and reviewing data analysis; and leading business analytics teams.  He is currently a Business Analytics Leader at Blue Sigma Analytics.  He provides services for everything from strategic consulting for business analytics to data analysis and data management.  His services are reflected by his book, workshops, and presentations.

He designed 'A Practitioner’s Guide to Business Analytics' (McGraw-Hill, 2013) (https://tinyurl.com/jx8rcru) to be the foremost reference on how corporations can better implement business analytics and in this era of Big Data and the Internet of Things.  He discusses strategic topics, including culture, organization, planning, and leadership for business analytics, in Chapters 1-6 and in Day I of his workshop.  For tactics, he discusses Statistical Qualifications, Diagnostics, and Review; and Data Collection, Software, and Management in Chapters 7-12 and during Day II of his workshop.  He previously contributed to the Encyclopedia for Research Design and writes blogs (including a series on Statistical Denial), case studies (AIG, AstraZeneca, big pharma, Google Flu Trends, et al.), and articles (two in Analytics Magazine). 

Specialties: Leading quants; addressing Big Data; making and supporting analytics-based decisions; performing statistical review; evaluating datasets and software needs;and re-organizing analytics teams.

Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.