Big Data is what underpins the analytics of most massive companies. With data flow running into the petabytes, analytics should be having a field day, leading to insights about all kinds of things that human analysts couldn’t predict. Yet somehow we’re not seeing any of these groundbreaking insights. Instead, we’re seeing more and more companies folding up their projects or starting new ones while leaving the old ones unresolved. ORMS Today mentions that the failure rate of large projects, according to Gartner, is around 85%. What seems to be the stumbling block that big data analytics can’t seem to overcome?
Bad Data Leads To Bad Results
Ever since the early days of computing, professionals have dealt with a phenomenon known as garbage in, garbage out, or GIGO. Tech Terms defines GIGO as a situation where bad data going into a program will result in erroneous or bad output. So, for example, if you’re using a spreadsheet to give you business insights, it can only offer you insights based on the data it has. if that data is bad, then the insights it gives will be nonsensical. This fact is the crux of the current problem with Big Data Analytics – the data just isn’t as good as it should be.
That’s not to say that Big Data doesn’t have a robust system of scrubbing. Many companies that collect Big Data sets go to great lengths to ensure that the data they retrieve is of the highest quality. Unfortunately, these methods aren’t perfect, and while the rudimentary filtering gets rid of the most glaring errors, it doesn’t deal with erroneous reports from the source. Consumers today have learned how to game data collection systems and the bad data that resides in these sets isn’t due to input error, but rather humans sidestepping the data-collection process.
Technology Largely Ignores Bad Data
Data analytics has advanced by leaps and bounds in recent years, with companies pumping millions of dollars into perfecting their algorithms. Data scientists are among the most well-paid programmers in the world, with their skills going towards making more efficient algorithms for processing incoming data streams. The problem is, the data in these streams may not be trustworthy. Without a filter to separate the bad data from the good data, the system will fail to deliver meaningful insights. As smart as these algorithms can be, they can only work within the confines of their parameters and with the data they’re presented.
Humans Are Problematic
The confines of separating garbage data from real data relies on where that data comes from. Big Data combined with AI has had a moderate amount of success in streamlining and dealing with data that conforms to one particular subject. AI Chatbots and customer relationships management software, for example, relies on predicting the outcomes of certain, finite things. Thus, by limiting their scope into something that’s more easily defined, they ask less of the analytics engine, and ensure that the data it’s using falls within a set of strictly defined parameters.
As mentioned before, human beings display a penchant for sidestepping the process of data collection. While this benefits them, algorithms (not having any idea of intent) can’t factor in this type of behavior. The result is a system that collects and processes data blindly. By trying to tap into Big Data to determine human behavior, data science has inadvertently modified that behavior against itself. It’s a testament to how adaptable human beings can be, and a lesson for Big Data analysts that insights are only as good as the data its based on.
When it comes to dealing with human beings, however, this consideration is even more important. While an analytics engine could theoretically look at human behaviors and predict outcomes, it’s never as simple as following a handful of parameters and modeling them accordingly. When we combine the fact that there may be myriad parameters to keep track of, and no way to guarantee the veracity of data, it becomes clear why these analytics projects fail. The truth is that they’re trying to model something too complex for analytics to comprehend, at least at this time.
Bad Data Leads to Bad Analytics
The essential takeaway from this understanding is that not all data is created equal. Yes, Big Data means endless streams of data that can be used to figure many details about a situation out, but once you introduce human beings into that data set, it becomes a lot more difficult to make headway into what those insights really mean. For small, limited topics, Big Data Analytics can be useful. However, as the dataset becomes more diverse and the parameters are linked to others which the analytics engine can’t see, the insights become less impactful. It’s a classic case of GIGO, and the only way to cure it is to increase the quality of the data that the analytics engine has access to.