Big Data ”large amounts of information that inundates businesses at high-speed and in a variety of formats ”is increasingly being processed and stored for analysis by companies. Businesses understand that they can leverage technology to derive crucial insights from this data. Such insights can drive down costs, streamline operations, and lead to more informed decisions.
This article defines a concept which is very relevant to Big Data; namely, tiered storage. After you finish reading, you’ll understand what tiered storage is, its advantages, and its applicability for Big Data.
Tiered Storage Defined
Tiered storage assigns data to two or more different types of storage media depending on the business value of that data. Storage tiering is an aspect of information management that deals with getting the most from business data at the lowest possible cost. The value of data often changes over time, and there are quite a few determining factors that dictate this value, such as how often you access the data, what capacity you need to support its uses and special circumstances that require its quick retrieval, such as an account audit or financial report.
Tiered storage uses a range of modern storage technologies, including solid state, hard disk, cloud storage, and tape storage, and the data that ends up in each type of storage depends on how you classify it and how many tiers you use. The complexities involved in classifying data value and manually moving data between different storage media has led to the emergence of automated storage tiering services that can apply policy-based rules to move data between different tiers automatically. The main vendors in automated storage tiering include IBM, NetApp, and EMC.
Advantages of Tiered Storage
-
Cost-effective: tiered storage drives huge reductions in storage costs compared to untiered storage. Using a single type of storage for all data is a waste of money for most businesses.
-
Operationally Efficient: With tiered storage, only the data that serves important business functions ends up in high-performance storage media, such as solid-state drives while archival data with low value ends up in tape storage or low-cost cloud storage services.
-
Flexibility: You have the flexibility, particularly with automated tiering, to move data between different storage media as its value changes.
Tiered Storage and Big Data
Distributed processing frameworks such as Hadoop are used for data processing and storage for Big Data applications. Hadoop’s massive storage capabilities come from its clustering architecture wherein data is distributed and stored in a network of multiple computing resources. You can either set up Hadoop on-premise in your data center or the cloud.
As large datasets enter Hadoop clusters, parts of the data are stored on individual machines or nodes in a cluster. In the initial stages after its creation, this data is frequently accessed by different teams who attempt to derive insights from it.
However, the frequency of access, as with many types of data, tends to decline over time. Hadoop supports data tiering, and by splitting your cluster into different tiers based on the frequency of use, you can dramatically reduce the costs of storing this data. The typical approach is a multi-temperature one that splits big data as follows:
-
Hot: data requiring high-frequency access for real-time analytics, reporting, or ad-hoc queries. This data resides on resources with maximum computing power.
-
Warm: less frequently accessed data that is still accessed enough to warrant residing on disks or even SSD storage.
-
Cold: archival data which is rarely accessed but which an enterprise wants to retain for compliance or once-off queries. This data resides on storage with minimal processing power available to it.
The reductions in cost come from splitting up the data so that the least frequently used data resides on nodes that have minimal computing power because storage without computing power is much cheaper. Data can flow between tiers via Hadoop tools such as Mover if it becomes more or less active, allowing for greater efficiency in Big Data storage.
Categorizing Big data into different storage tiers based on its frequency of usage is a good starting point. To further streamline your storage requirements, you don’t need to wait for data to age before shifting it to low-cost storage tiers.
Out of the mountains of Big Data gathered by businesses every day, only a small proportion ends up being useful, so there is the potential to classify and tier your Big Data as soon as it enters your clusters, leading to even greater efficiencies.
Wrap up
Storage tiering has great potential in a world where businesses are under pressure to gain useful insights from the large swathes of data they collect on a regular basis; data that will only continue to grow in volume and velocity. Tiered storage brings cost optimizations to the table that can ensure organizations achieve the right balance between performance, capacity, and cost on their big data clusters.