Big data is an enormously popular concept and at the first glance, it does sound like a silver bullet. Whatever your business deals with, whatever your problem is, for that matter, all you have to do is collect the sufficient amount of data, plug it all in the computer and wait for insights. The long-awaited answers will pour out and tell you how to increase your sales, what demographics to target, and how to make your employees more loyal and efficient.
However, the reality is more complicated than that. Often information is incomplete, therefore misleading. Moreover, data has a bad habit to change constantly. People move homes and change their jobs and phone numbers, start families, gain or lose weight, dye hair, start or quit shaving, develop new medical conditions and eating habits. Your data is decaying by the hour. It is important that you do regular revision and updates of your databases. Cleaning up the dirty data is an important and often overlooked task, necessary to prevent costly mistakes.
What do you mean dirty ?
Let’s make it clear: there are no clean data sets. The life is complicated, messy and full of white noise and irrelevant facts. Even the most tidy and organized database reflects this irregularity to some extent. The question is how much of the garbage you have in your data and how big is a margin of error in each particular research you draw from it.
According to Technopedia, there are the following subsets of dirty data:
- Misleading data
- Duplicate data
- Incorrect data
- Inaccurate data
- Non-integrated data
- Data that violates business rules
- Data without a generalized formatting
- Incorrectly punctuated or spelled data
When it comes to marketing research and health care duplicate and inaccurate (more specifically, incomplete data) are the main woes. In medical care industry, each practitioner has their own way of putting things into records “ abbreviations, keywords, detailed descriptions. Sometimes they neglect to fill out a particular field in the form because they deem it irrelevant. Due to all this, it is virtually impossible to slice and dice the big data obtained this way without previous revision and unification. Before pharmaceutical companies or researchers can use it, this data must be cleaned-up.
Where does the dirt come from?
Currently, there are four main sources of data collection: existing datasets, social media, smartphones, and wearables. The latter usually go with the designated apps that ask users for additional information and collect relatively reliable self-reported data. However, users do make mistakes or simply neglect to enter all the details.
Existing datasets, however well kept, need updates. Moreover, to elicit additional insights, researchers merge databanks obtained from different sources. Sometimes, improper data merge leads to duplicate data occurrence. Duplicates may also appear due to repeated submissions. The simplest example of duplicate data sets is duplicate files on your computer. You save the files in the different directories and system thinks they are two (three, four) completely different files. This problem, however, is easy to solve with Macflypro or similar cleaning software (if only everything was so easy in the big data science!)
Another source of incorrect data is typos, to put it simply. Data is incorrect when field values are created outside of the valid range of values. For instance, the value in a month field should range from 1 to 12. If someone put 13 there, the value is void. You can assume they meant 12 or 03, or they put the date in the wrong field, but there is no way of knowing.
Sometimes mistakes are intentional “ users simply do not want to provide a valid information, they misspell their names, invent email addresses and mess up the numbers of their area codes or phone numbers to remain anonymous or avoid unsolicited calls.
What can we do?
When it comes to customer database, complicated algorithms must be employed to analyze it for duplicates and other mistakes. The best cure, in this case, is prevention, i.e. improved procedures of data collection, forms and records with more standardized fields.
Embedded analytics works in a similar way. When it is built in the business operations, it allows you to identify error/warning conditions, respond rapidly and eliminate dirty data before it is added to your database.
You should also make origins and history of data visible and transparent “ this way you will be able to trace back every mistake. When put in context, such mistakes may turn out to be pieces of valuable additional information. Alternatively, you will identify an unreliable source of information and thus will be able to eliminate it from further data collection process.
Overall, the completeness, validity, and consistency of your data depend on the methodology you use. Each business needs to develop and customize such methodology for its particular needs.
Data analysts generally look for irregularities, something that is out of the norm or simply seems off . This part of the job is usually automated due to the vast volumes of input data. Thanks to machine learning, software becomes more and more competent in detecting what is right and what is wrong, but still, this is the first crude selection. The more detailed and accurate interpretation must be left to humans.
Bottom line
However, cleaning data is just a first step in the long process of interpretation. After all, we, not the machines draw conclusions from the big data. Computers are good at collecting, storing, crunching, sorting through millions of pages, but they will never provide a complete answer to your problem, no matter how clean your dataset is. To get the result, you must go through a grueling and time-consuming process of evaluation, analysis, and interpretation.
Alternatively, look at the dirty data as an opportunity “ there is always a way to learn from your mistakes. Google spelling suggestions are a good example of how dirty data can lead to improvement. It took years of collecting misspelled queries to make a system that runs so smoothly now. Ultimately, there is no dirty data, there is just not enough data to make the sense of it. So keep collecting “ that’s the only way forward.