Posted in

Build Production Ready clusters for Big Data workloads

CDH is an open-source distribution of Cloudera, which includes Apache Hadoop built to meet the Enterprise demands of customers. It is a modern platform to establish production-ready clusters for multiple environments.

In this post, we illustrate some best practices and guidelines for the next part of the implementation process.

Network Considerations

HOSTNAME RESOLUTION, DNS AND FQDNS

Cluster hosts must have a working network name resolution system and correctly formatted /etc/hosts file. All cluster hosts must have properly configured forward and reverse host resolution through DNS. The /etc/hosts files must:

  1. Contain consistent information about hostnames and IP addresses across all hosts
  2. Not contain uppercase hostnames
  3. Not contain duplicate IP addresses

You can verify that forward and reverse lookups are configured correctly using the Linux host command.

$ host `hostname`

bp101.cloudera.com has address 10.20.195.121

Java Requirements

Hadoop runs on top of JVM. Only 64 bit JDKs are supported.

  • JDK 7
  • JDK 8

Storage and Disks Considerations

The entire Hadoop ecosystem was created by keeping in mind the JBOD configuration. HDFS is an immutable filesystem that was designed for large files. This goal plays well with stand-alone SATA drives, as they get the best performance with sequential reads. RAID is used to add redundancy to an existing system, HDFS already has that built-in.

Cloudera recommends that only a single directory be used if the underlying disks are configured as RAID, or two directories on different disks if the disks are mounted as JBOD.

Operating System Prerequisites

Before deploying the Hadoop cluster, it’s necessary to check the following:

SELINUX

Security-Enhanced Linux (SELinux) allows you to set Access control using policies. However, Selinux kernel can interrupt the installation of Hadoop, so most of the customers run with SELinux disabled.

setenforce 1

IPTABLES

Cloudera recommends disabling internal firewalls on the cluster, at least until the cluster is up and running. As the firewall may interuppt in the passwordless login for Hadoop.

$ service iptables stop

$ service ip6tables stop

$ chkconfig iptables off

$ chkconfig ip6tables off

IPV6

Hadoop 2 does not support IPv6. IPv6 configurations should be removed, and IPv6-related services should be stopped.

swappiness

$ sysctl vm.swappiness = 1

$ echo “vm.swappiness = 1” >> /etc/sysctl.conf

TRANSPARENT HUGE PAGES (THP)

Transparent Huge Page compaction interacts poorly with Hadoop workloads and can seriously degrade performance. It’s recommended to disabling THP to avoid Defragmentation.

$ echo ‘never’ > defrag_file_pathname

NTP

CDH requires that you configure a Network Time Protocol (NTP) service on each host of your cluster. Most operating systems include the ntpd service for time synchronization.

$ service ntpd start chkconfig ntpd on

Required Databases

Cloudera Manager uses various databases to store information about the Cloudera Manager configuration, and other information such as the health of the system, or task progress and provide the metric.

The following components all require databases: Cloudera Manager Server, Oozie Server, Sqoop Server, Activity Monitor, Reports Manager, Hive Metastore Server, Hue Server, Sentry Server, Cloudera Navigator Audit Server, and Cloudera Navigator Metadata Server.
Recommended And supported databases:

  1. MariaDB
  2. Mysql
  3. Oracle DB
  4. PostgreSQL

Leveraging appropriate Services

CDH provides services depending on the Applications such as data science applications, ETL applications, or data analytics services.

Hadoop Essentials

HDFS, YARN, MapReduce 2, ZooKeeper, Oozie, Hive, and Hue

Data Engineering

HDFS, YARN, MapReduce 2, ZooKeeper, Oozie, Hive, Hue, and Spark

Data Warehouse

HDFS, YARN, MapReduce 2, ZooKeeper, Oozie, Hive, Hue, and Impala

Operational Database

HDFS, YARN, MapReduce 2, ZooKeeper, Oozie, Hive, Hue, and HBase

All Services (Cloudera Enterprise Data Hub)

HDFS, YARN, MapReduce 2, ZooKeeper, Oozie, Hive, Hue, HBase, Impala, Solr, Spark, and Key-Value Store Indexer

Custom Services

Choose your own customized services. Flume can be added after your initial cluster has been set up.

Service and Host Placement

Master Host: Runs the Hadoop master daemons: NameNode, Standby NameNode, YARN Resource Manager and History Server, the HBase Master daemon, Sentry server, and the Impala StateStore Server and Catalog Server.

Worker Host: Runs the HDFS DataNode, YARN NodeManager, HBase RegionServer, Impala impala.

Utility Host: Runs Cloudera Manager and the Cloudera Management Services.

Edge Host: Contains all client-facing configurations and services, gateway configurations for HDFS GW, YARN GW, Impala GW, Hive GW, and Hbase GW.

Client Configuration Files

Client configuration files are generated automatically by Cloudera Manager based on the services and roles you have installed.

Client configuration files are deployed on any host that is a client for service on the cluster.

  • hadoop-env.sh
  • core-site.xml
  • hdfs-site.xml
  • mapred-site.xml
  • yarn-site.xml
  • log4j.properties

After the installation and deployment of cluster, you can view the Cloudera Management Admin Console.

Conclusion

 

This is how Cloudera has made it easy for you to deploy, configure, and manage CDH cluster using Cloudera distribution of Hadoop.

I am a full-time guys and a part-time blogger. Daniel Jacob is a globally writer for a big data, artificial intelligence, machine learning, data analytics, python and other emergency technologies. He holds a bachelor of Technology in New York Institute Technology.

Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.