Posted in

Five Patterns of Big Data Integration

Intel Data Center GPU codenamed Crescent Island architectural slide showcasing Xe3P AI optimized GPU IP, up to 480GB LPDDR5x memory capacity, and a 350W air-cooled PCIe form factor.
Intel’s Crescent Island GPU targets AI inference economics by prioritizing massive LPDDR5x memory capacity over costly HBM architectures within a practical 350W air-cooled design.

As reliance on Hadoop and Spark grows for data management, processing and analytics, data integration strategies should evolve to exploit big data platforms in support of digital business, Internet of Things (IoT) and analytics use cases. While Hadoop is used for batch data processing, Spark supports low-latency processing. Integration leaders should understand the various patterns of integration described below, and align use cases with vendor offering.

1) Native ETL and ELT in Apache Hadoop & Spark platforms  data integration occurring natively in the Hadoop & Spark platforms without leveraging data integration tools, but rather using native tools (such as Pig, Hive, MapReduce). Skillset scarcity is a potential concern with this pattern.

2) Data integration offering specific to Hadoop & Spark platforms (distinct from traditional data integration offering)  incumbent Vendors providing a dedicated big data integration offering (distinct from their traditional data integration offering) that runs in Hadoop & Spark platforms. A separate offering potentially avoids disruption to traditional data integration workflows.

3) The same offering for both traditional and big data integration needs  vendors leveraging the same offering for traditional data and big data integration.

  • Incumbent vendors evolving their traditional data integration product to support big data integration.
  • Emerging vendors providing data integration products that support both traditional and big data integration.

4) Data pipelines in Hadoop & Spark platforms  vendors providing end-to-end data management solution (ingestion, organization, transformation, enrichment, and quality) including integration.

  • Vendors providing end-to-end analytics solution (ingestion, organization, transformation, enrichment, and analytics) including integration.
  • Vendors providing a framework for building and deploying data applications including integration

5) Self-service data preparation using Hadoop & Spark platforms  self-service data preparation offerings using Hadoop & Spark platforms to support the processing requirements for data preparation.

Lakshmi Randall is an industry analyst and strategist providing independent advisory services in information management and analytics domains. She has over 19 years’ experience working for industry leaders such as Microsoft. She is a strong advocate of data preparation, and published research at Gartner on its importance to analytics. Among her clients, she has a reputation as an informative, trusted adviser who never comprises her objectivity. 

 

Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.