Without effective and comprehensive validation, a data warehouse becomes a data swamp.
With the accelerating adoption of Snowflake as the cloud data warehouse of choice, the need for autonomously validating data has become critical.
While existing data quality solutions provide the ability to validate Snowflake data, these solutions rely on rule-based approach that are not scalable for 100s of data assets and often prone to rules coverage issues. More importantly, these solutions do provide an easy way to access audit trail of results.
According to a 2021 study by Boston Consulting Group, data quality is lagging in most companies.
Current Approach and Challenges
The current focus in Snowflake Data warehouse projects is on data ingestion, the process of moving data from multiple data sources (often of different formats) into a single destination. After data ingestion, data is used and analyzed by business stakeholders – which is where data errors/issues begin to surface. As a result, business confidence in the data hosted in Snowflake reduces. Our research estimates that an average of 20-30% of any analytics and reporting projects in snowflake is spent identifying and fixing data issues. In extreme cases, the project can get abandoned entirely.
Current data validation tools are designed to establish data quality rules for one table at a time – as a result there are significant cost issues in implementing these solutions for 100s of tables. Table wise focus often leads to incomplete set of rules or often not implementing any rules for certain tables resulting in unmitigated risks.
Despite significant investments in data quality solutions, most organizations are not able to ensure quality in their data assets because of following challenges:
High Cost of Implementation: Existing Data Quality solutions rely on a rule-based approach. As a result, implementation effort is linearly proportional to the number of tables in Snowflake. Maintaining thousands of implemented rules as the data evolves adds to the total cost of ownership
Architectural Limitations: Many of the existing tools are not architected to validate billions of records that some of the Snowflake tables may contain. In addition, Data needs to be moved from Snowflake to the Data Quality solution, resulting in latency as well as significant security risks.
Knowledge Gap: Data quality analysts are often not familiar with the data assets. In order to create data quality rules, they need to consult subject matter experts extensively. In the Snowflake Data Cloud, as organizations share datasets with each other – Data Quality analysts may not have access to the subject matter experts from another organization.
What is DataBuck?
DataBuck is an autonomous “Powered by Snowflake” data validation solution for Snowflake. It establishes data fingerprint and an objective data trust score for each data asset (Schema, Tables, Columns) presents in Snowflake using its ML capabilities. Trust in data will no longer be a popularity contest. No need to have individuals give their subjective opinion on the health of a table/file. All stakeholders can universally understand the objective Data Trust Score.
More specifically, it leverages machine learning to measure the data trust score through the lens of standardized data quality dimensions as shown below:
1. Freshness – determine if the data has arrived before the next step of the process
2. Completeness – determine the completeness of contextually important fields. Contextually important fields should be identified using various mathematical and or machine learning techniques.
3. Conformity – determine conformity to a pattern, length, format of contextually important fields.
4. Uniqueness – determine the uniqueness of the individual records.
5. Drift – determine the drift of the key categorical and continuous fields from the historical information
6. Anomaly – determine volume and value anomaly of critical columns
DataBuck can auto-trigger Data Trust Score as soon as new data lands in a Snowflake table or can be scheduled to run at a specific time or as part of the data pipeline.
How does DataBuck work?
The user provides Snowflake connection information along with the database details and triggers the continuous data validation process. Once the data validation process is activated, DataBuck sends its ML engine to snowflake to analyze the data and identify data quality issues. Summary results are then presented to the user through the web console. At no point in this process, the user needs to write rules or move data out of snowflake.

Setting it up in 60 Seconds.
As shown below, the user follows the following process to setup :
- Provides database and schema name for which data validation needs to be done

2. Indicates whether continuous data validation needs to be performed or not
3. Trigger the data validation process by clicking the health check button.
DataBuck results:
It can validate a snowflake database regardless of the number of tables and size of each individual table. It provides the following results:
- Data Quality of a Schema Overtime:

2. Summary Data Quality Results of Each Table

3. Detailed Data Quality Results of Each Table

4. Detailed Data Profile of Each Table

5. Discovered Data Quality Rules for Each Table

Summary
Data is the most valuable asset for modern organizations. Current approaches for validating data, in particular SNOWFLAKE, are full of operational challenges leading to trust deficiency, time-consuming, and costly methods for fixing data errors. There is an urgent need to adopt a standardized autonomous approach for validating the SNOWFLAKE data to prevent data warehouse from becoming a data swamp.
DataBuck provides a secure and scalable approach to validate snowflake data in an ongoing manner. All it takes is a single click and you can validate hundreds of your Snowflake tables.