Posted in

Preparing for Disaster: The Importance of Big Data Disaster Recovery

Disaster recovery encompasses procedures, policies, and techniques aimed at ensuring the swift and smooth recovery of organizational data and infrastructure after a disaster, whether human or natural-induced.

The best way to gain a true understanding of the importance of disaster recovery is to read about some real-life disaster recovery stories in which the lack of clear strategy and tests caused problems:

  • In May 2017, an IT worker accidentally switched off British Airways’ power supply, causing a prolonged outage. The outage resulted in the cancellation of at least 800 flights and compensation costs of over $70 million.

  • Pixar almost lost all of its files for hit movie Toy Story 2 when an employee accidentally deleted them, and the company realized its backups hadn’t functioned for two months because they hadn’t been properly tested. The only reason the project was retrieved was that an employee had taken home a copy of the movie’s files completely by chance the previous week.

Cloud computing has made disaster recovery more affordable and accessible for business with the availability of disaster recovery-as-a-service. Major vendors include Azure Site Recovery, IBM Resilience Services, and Veritas. Other service providers, such as N2WS, recognize that cloud environments can also be vulnerable, and they offer a service that backs up cloud-based workloads with disaster recovery solutions for AWS and more. After all, Amazon Web services (AWS) and other cloud service providers have experienced major outages before.

Find out in this post why Big Data disaster recovery is becoming increasingly important for businesses.

How Disaster Recovery Is Relevant For Big Data

Big Data analytics tools, software, and workloads are now mission-critical for enterprises looking to derive insights and a competitive edge from the large stores of information that inundates their systems daily. As far back as 2013, 70 percent of IT executives deemed Big Data mission-critical.

Businesses typically use a range of tools to analyze and process Big Data, such as Apache Spark, Hadoop, Kafka, and Apache Storm. The mission-critical nature of Big Data workloads shows that it’s vital to include Big Data environments in any disaster recovery plan.

Since what qualifies as Big Data is often high-velocity data gathered in real-time (IoT sensors, web clicks) the speedy restoration of environments after a disaster takes on arguably even greater importance with this type of data. Lose an hour of streaming information and you might be ok, but lose half a day or more, and you’ll likely miss out on some critical business insights.

Disaster recovery provides insurance against unexpected outages; insurance that is invaluable for the data-driven businesses that operate in the modern world.

Big Data Disaster Recovery Tips

Establish and Agree Upon An RPO

In the context of disaster recovery, the recovery point objective (RPO) is the interval that equates to the maximum acceptable amount of data loss after an unplanned outage occurs. The RPO is so crucial because it determines the frequency at which you need to back up your data.

There is often confusion about what the RPO actually means, and this confusion mostly arises due to quite a variation in definitions. To clarify RPO, ask yourself how many hours of data loss as a result of Big Data systems going down can you afford. If your RPO is two hours of data, then you perform backups every two hours.

Make sure all important stakeholders, including IT executives, management, IT operators, etc., are all clear on what the RPO means and that everyone agrees on the same figure.

Conduct Regular Disaster Recovery Tests

You can’t have confidence in any disaster recovery plan without conducting regular testing of it. The Pixar story previously mentioned in this article is a case in point ”the team neglected to test their backup procedures, which form a pivotal part of any disaster recovery strategy, and they almost paid a high price.

Only you can decide a testing frequency that works for your business, but a safe approach is to test at least semi-annually for newly implemented disaster recovery plans and then yearly going forward.

The tests you conduct should verify that your disaster recovery procedures and strategies can restore Big Data workloads within pre-defined RPO and RTO (recovery time objective). Your RTO is the target maximum time for resuming your Big Data workloads. The mission-critical nature of Big Data points to a short RTO, and the shorter the RTO, the more resources required to achieve it.

Back up Off-Premise

Whether in low-cost cloud storage services or in a second physical location, it’s always prudent to back up any data to an off-premise location. While backup is not synonymous with disaster recovery, creating data backups at specified intervals is a huge part of any successful disaster recovery strategy.

For Big Data, cloud-based backup is probably the best option because it is cheap and easy to back up your data to the cloud, particularly batch data, which is large and static.  

Seek Out a Converged Data Platform

Backups alone aren’t enough for disaster recovery, and it makes sense to seek out some kind of converged data platform for your Big Data disaster recovery. On a converged platform, you can manage multiple Big Data clusters across several locations and infrastructure types (cloud/on-premise) independent of service provider, ensuring data remains consistent and up-to-date between all clusters.

For example, IBM Big Replicate is one such platform that replicates data across multiple clusters and unifying Hadoop clusters running on different vendor distributions and versions. The active replication of IBM Big Replicate drives both RTO and RPO to near-zero durations.  

Conclusion

The first step in ensuring you’re prepared for disaster with Big Data workloads is to recognize the mission-critical nature of the insights these workloads can deliver.

Adapt your disaster recovery plan to include a section on Big Data, and make sure to implement some of the tips in this article to increase the chances of returning Big Data environments to an acceptable service level as quickly as possible after a disaster.

I'm a technical writer and editor with over 10 years' experience writing technical articles and documentation for various audiences, including technical on-site content, software documentation, and dev guides. I specialize in big data analytics, computer/network security, middleware, software development and APIs. And I love coffee!

I've published my work on major publications such as DZone or Wordtracker, and I'm also a volunteer writer for some universities, where I write about data science, big data, data warehousing, and related topics. Here are some of my recent articles:

Soft Computing vs. Hard Computing (UoPeople)
An Overview of Amazon Redshift (DZone)
Managing Telehealth’s Big Data with Data Warehousing (Arizona University)
Facial Features, Illnesses, and Computer Vision (Pompeu Fabra University)
Tools for a Deeper Understanding of User Data (Wordtracker)

Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.