CDH is an open-source distribution of Cloudera, which includes Apache Hadoop built to meet the Enterprise demands of customers. It is a modern platform to establish production-ready clusters for multiple environments.
In this post, we illustrate some best practices and guidelines for the next part of the implementation process.
Network Considerations
HOSTNAME RESOLUTION, DNS AND FQDNS
Cluster hosts must have a working network name resolution system and correctly formatted /etc/hosts file. All cluster hosts must have properly configured forward and reverse host resolution through DNS. The /etc/hosts files must:
- Contain consistent information about hostnames and IP addresses across all hosts
- Not contain uppercase hostnames
- Not contain duplicate IP addresses
You can verify that forward and reverse lookups are configured correctly using the Linux host command.
$ host `hostname`
bp101.cloudera.com has address 10.20.195.121
Java Requirements
Hadoop runs on top of JVM. Only 64 bit JDKs are supported.
- JDK 7
- JDK 8
Storage and Disks Considerations
The entire Hadoop ecosystem was created by keeping in mind the JBOD configuration. HDFS is an immutable filesystem that was designed for large files. This goal plays well with stand-alone SATA drives, as they get the best performance with sequential reads. RAID is used to add redundancy to an existing system, HDFS already has that built-in.
Cloudera recommends that only a single directory be used if the underlying disks are configured as RAID, or two directories on different disks if the disks are mounted as JBOD.
Operating System Prerequisites
Before deploying the Hadoop cluster, it’s necessary to check the following:
SELINUX
Security-Enhanced Linux (SELinux) allows you to set Access control using policies. However, Selinux kernel can interrupt the installation of Hadoop, so most of the customers run with SELinux disabled.
setenforce 1
IPTABLES
Cloudera recommends disabling internal firewalls on the cluster, at least until the cluster is up and running. As the firewall may interuppt in the passwordless login for Hadoop.
$ service iptables stop
$ service ip6tables stop
$ chkconfig iptables off
$ chkconfig ip6tables off
IPV6
Hadoop 2 does not support IPv6. IPv6 configurations should be removed, and IPv6-related services should be stopped.
swappiness
$ sysctl vm.swappiness = 1
$ echo “vm.swappiness = 1” >> /etc/sysctl.conf
TRANSPARENT HUGE PAGES (THP)
Transparent Huge Page compaction interacts poorly with Hadoop workloads and can seriously degrade performance. It’s recommended to disabling THP to avoid Defragmentation.
$ echo ‘never’ > defrag_file_pathname
NTP
CDH requires that you configure a Network Time Protocol (NTP) service on each host of your cluster. Most operating systems include the ntpd service for time synchronization.
$ service ntpd start chkconfig ntpd on
Required Databases
Cloudera Manager uses various databases to store information about the Cloudera Manager configuration, and other information such as the health of the system, or task progress and provide the metric.
The following components all require databases: Cloudera Manager Server, Oozie Server, Sqoop Server, Activity Monitor, Reports Manager, Hive Metastore Server, Hue Server, Sentry Server, Cloudera Navigator Audit Server, and Cloudera Navigator Metadata Server.
Recommended And supported databases:
- MariaDB
- Mysql
- Oracle DB
- PostgreSQL
Leveraging appropriate Services
CDH provides services depending on the Applications such as data science applications, ETL applications, or data analytics services.
Hadoop Essentials
HDFS, YARN, MapReduce 2, ZooKeeper, Oozie, Hive, and Hue
Data Engineering
HDFS, YARN, MapReduce 2, ZooKeeper, Oozie, Hive, Hue, and Spark
Data Warehouse
HDFS, YARN, MapReduce 2, ZooKeeper, Oozie, Hive, Hue, and Impala
Operational Database
HDFS, YARN, MapReduce 2, ZooKeeper, Oozie, Hive, Hue, and HBase
All Services (Cloudera Enterprise Data Hub)
HDFS, YARN, MapReduce 2, ZooKeeper, Oozie, Hive, Hue, HBase, Impala, Solr, Spark, and Key-Value Store Indexer
Custom Services
Choose your own customized services. Flume can be added after your initial cluster has been set up.
Service and Host Placement
Master Host: Runs the Hadoop master daemons: NameNode, Standby NameNode, YARN Resource Manager and History Server, the HBase Master daemon, Sentry server, and the Impala StateStore Server and Catalog Server.
Worker Host: Runs the HDFS DataNode, YARN NodeManager, HBase RegionServer, Impala impala.
Utility Host: Runs Cloudera Manager and the Cloudera Management Services.
Edge Host: Contains all client-facing configurations and services, gateway configurations for HDFS GW, YARN GW, Impala GW, Hive GW, and Hbase GW.
Client Configuration Files
Client configuration files are generated automatically by Cloudera Manager based on the services and roles you have installed.
Client configuration files are deployed on any host that is a client for service on the cluster.
- hadoop-env.sh
- core-site.xml
- hdfs-site.xml
- mapred-site.xml
- yarn-site.xml
- log4j.properties
After the installation and deployment of cluster, you can view the Cloudera Management Admin Console.
Conclusion
This is how Cloudera has made it easy for you to deploy, configure, and manage CDH cluster using Cloudera distribution of Hadoop.