Posted in

Survey Shows Enterprises Struggling with Bad Data

Last week we announced the results of a survey of over 300 enterprise data professionals conducted by Dimensional Research and sponsored by StreamSets. We were trying to understand the markets state of play for managing their big data flows. What we discovered was that there is an alarming issue at hand: companies are struggling to detect and keep bad data out of their stores.

There is a Bad Data Problem Within Big Data

When we asked data pros about their challenges with big data flows, the most-cited issue was ensuring the quality of the data in terms of accuracy, completeness and consistency, getting votes from over of respondents. Security and operations were also listed by more than half. The fact that quality was rated as a more common challenge than even security and compliance is quite telling, as you usually can count on security to be voted the #1 challenge for most IT domains.

Challenges Big DataThe painful reality of this challenge was hammered home by the number of people who admitted to flowing bad data into their stores (87%) or knowingly having bad data in their stores (74%). On an equally disturbing note, nearly one in 8 respondents (12%), answered I dont know to the question about bad data in their stores, which may point to an issue around data governance.

Data Drift Plus Hand Coding Create Quality Issues

At StreamSets we believe that the quality issue within big data flows is related to data corrosion and data loss caused by data drift. Data drift is defined as unexpected, unannounced and unending changes to schema and semantics that occur at the data source and usually go undetected as they flow into data stores or to analytic applications. In the survey, a dramatic majority of respondents acknowledged that data drift creates a significant impact on their data flow operations such as pipeline failures, slowdowns or data corruption.

Part of the reason for this high impact from data drift may be the continued reliance on hand-coded data pipelines. Low-level coded solutions tends to be brittle in the face of schema changes and usually are not instrumented to monitor, detect or deal with changing schema or semantics (opaqueness). Nearly 4 in 5 respondents still use hand-coding. As you can see ETL tools are also heavily used and they too suffer from being brittle in the face of change and opaque when it comes to operational visibility.

You Cant Address What You Cant Detect

Lastly, we also surveyed on abilities and desires related to different areas of operational management. Given this preponderance of bad data caused in part by data drift and hand coding, its not surprising that detecting changes to data while it is in motion is the domain where enterprises felt weakest, and where there was the biggest gap between current capabilities and preferred state. Only 34% graded themselves as good or excellent at detecting data divergence; but more than twice that percentage (69%) considered detecting divergent data to be a valuable capability.

Rick is the VP of Marketing at StreamSets, a company that delivers performance management for data flows. He is a marketing leader with success across technology startup, mid-size and Fortune 500 companies. Rick specializes in the creation and execution of innovative marketing programs and best practices that build brands and drive demand. He has an MBA from Stanford University. 

Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.