Posted in

10 Things to Consider Before Diving Into the Hadoop Data Lake

A conceptual technical diagram illustrating the AWS and Google Cloud Interconnect partnership. It shows jigsaw puzzle pieces of AWS, Azure, Google Cloud, and Oracle, connected by a dedicated network pipe. Accompanying news headlines reference cloud competition and regulatory scrutiny. A pie chart shows cloud market share. Gauges track multicloud spend optimization.
Connecting Rivals: The AWS and Google Cloud agreement establishes a direct, private link between their networks, bypassing the public internet, but analysts suggest the move is more about defining multicloud networking standards than pure customer ease. (Visual: A representation of cloud interoperability versus market control.)

So youre finally ready to dive into the Hadoop Data Lake?  Wherever you are in your journey, we at Zaloni have develop a list of considerations that should be kept front-and-center along the way. To that end weve created a Data Lake Considerations Checklist. At a minimum, a checklist is a communication tool, and the one below can be used in your organization on your Data Lake journey to inform and engage stakeholders across the organization, from the CIO to the developers.

1. Business Benefit Priority List

What business value will the cluster generate? Identify what objectives are tied to the core/essential business vs. those that are value-add (e.g. for a bank, fraud detection vs. twitter sentiment analysis). Stating the business uses along with the priority will guide everything from cluster configuration (e.g. for YARN queue configuration) to ecosystem tool choices.

2. Architectural Oversight

What technologies, frameworks, and tools need to be evaluated based on use cases prioritized by business and considering skills in the organization? Architecture review processes and gating processes for new systems can help maintain the Data Lakes architectural integrity.

3. Security Strategy

What are the policies and procedures for accessing the data and cluster resources? Are there any business processes that are impacted? A security strategy, including procedures and protocols for handling breaches should be implemented.

4. I/O & Memory Model

Will the cluster be used for running statistical models or for ETL? Memory and I/O fundamentals should be reviewed and modeled out for future growth. Some Data Lakes serve as an MPP engine whereas some function as a large database, so it is important to understand the hardware and compute needs of your Data Lake.

5. Workforce Skill-set Evaluation

What skills exist within the firm that can be leveraged? Understanding your current strengths will guide the ecosystem tool selection in the near term and will help guide human resource planning.

6. Five-Year Vision

Understanding the five-year goal of the cluster will allow you to plan strategies for cluster management and code & metadata organization.

7. Data Governance Strategy

What policies (e.g. lifecycle management policies such as retention, archiving, purging), audit/review procedures and authorities/roles are required to meet the business needs?

8. Operations Plan

How are cluster resources, workloads and users managed? What are the procedures for making requests or investigating issues? Understand the controls of your Data Lake environment to guide how your processes and procedures can be leveraged or need to change.

9. Communication Plan

How are cluster-wide changes communicated? How are requests triaged? How does a developer onboarding ramp up quickly? By asking these questions, you may find that distribution lists and possibly, new roles and responsibilities, are required in the organization.

10. DR Plan

What are the plans should the data center become unavailable? A DR plan is part of most enterprise systems, but in the early phases of Hadoop adoption DR is often an afterthought. Run through disaster scenarios early in new use case implementations and update the DR plan as required.

Zaloni has helped large firms across many industries on their journey from initial Hadoop adoption to large enterprise-wide production cluster implementations. We hope this checklist will help bring your teams together to talk about and explore the Data Lake considerations that otherwise might have been overlooked or assumed.

Image: Andy Spearing

Craig Lukasik is Senior Solution Architect at Zaloni, Inc. Craig is highly experienced in the strategy, planning, analysis, architecture, design, deployment and operations of business solutions and infrastructure services. He has a wide range of solid, practical experience delivering solutions spanning a variety of business and technology domains, from high-speed derivatives trading to discovery bioinformatics. Craig is passionate about process improvement (a Lean Sigma Green Belt and MBA) and is experienced with Agile (Kanban and Scrum). Craig enjoys writing and has authored and edited articles and technical documentation. When he’s not doing data work, he enjoys spending time with his family, reading, cooking vegetarian food and training for the occasional marathon.

Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.