As we have seen, the acceptance of big data was highly reassuring in 2016. More and more organizations started adding it to their top line priority for storing, processing, and extracting actionable value out of data in all sizes and forms. In the first half of 2017, the demand for systems which support huge volume of unstructured and structured data is on the rise.
The top market demand is for the platforms which help database managers to secure and use big data along with empowering the end users to analyze that data. These systems are now maturing to operate right inside the enterprise IT systems by meeting all standards. Further, we will discuss some key components of the changing big data scenario to keep an eye on.
1. Big data is now instantly approachable
It is possible now to run machine learning and sentiment analysis by using tools like Hadoop, but the primary questions people may often ask would be “How quick is the interactive SQL?” The database, after all, is the conduit of business users who want Hadoop data for a faster and repeatable KPI dashboard and exploratory analysis.
The need for higher speed has also fueled adoption of much faster databases such as MemSQL and Exasol as well as Kudu, which is a Hadoop-based store. SQL-on-Hadoop engines like Hive LLAP, Apache Impala, Phoenix, Presto, and OLAP-on-Hadoop and Drill technologies like Jethro Data, AtScale, and Kyvos Insights also act as accelerators to blur the lines between traditional warehouses and the new world of big data.
2. Big data is not Hadoop anymore
During the previous years of development, we have seen various technologies rising and demising in the big data wave to fulfill the needs for Hadoop analytics. However, big enterprises with the heterogeneous environment may no longer want to adopt a single data source as Hadoop. The answers to their wider range of questions are buried at a variety of sources ranging from cloud warehouses to even individual systems. These may include structured or unstructured data from non-Hadoop resources as well. Even the relational databases like SQL Server 2016 have become big data ready.
In the year 2018, customers may have the demand to do analytics on all these diversified sets of data. The data-agnostic and source platforms may survive, while the others which are only purpose-built for Hadoop may fail to deploy in all cases. We can see the exit of Platfora as an early indicator of this shift.
3. Leveraging data lakes for the get-go to drive value
We can consider a data lake as a human-made reservoir. In this mode, you first dam the end to build a cluster and then you let it fill with data. Once you establish such a lake, you start to use this data for various purposes like predictive analytics, cyber security, ML and much more.
Up to this point, hydrating this lake has been the end in itself. It will change in 2017 as Hadoop’s business justification requirements get tightened. Organizations may further need an active use of this lake to get quicker answers for on-the-go decision making. They may carefully consider the business outcomes before investing high on the infrastructure. It tends to foster a stronger and much in-depth partnership between business, IT, and analytics. With this, self-service platforms may gain a deeper and stronger recognition as tools for harnessing big data assets.
Many focus on Salesforce training and discuss the integration of Salesforce with big data. However, as an enterprise database admin, if you want to adopt or strengthen these technologies, it is essential to learn these from the market leaders to be competent and reap the most desired results.
4. Diversity, not velocity or volume will drive the big data investments
Gartner defined big data with three key V’s as:
- Volume
- Velocity
- Variety
Even though these three V’s are steadily growing, “variety” is lately becoming the single biggest consideration as far as big data investments are concerned. This trend is expected to continue for the rest of 2017 and further as enterprises may try to integrate more sources and focus on the big data “long tail.”
We can also see that the data formats are multiplying day by day from schema-free JSON to relational nested types and further non-flat data like Parquet, Avro, and XML data formats. So, the analytics platforms may get further evaluated on the basis of their capacity to provide direct and live direct connectivity to all these diversified data sources and formats.
5. No more one-size-fits-all type frameworks
As we discussed earlier, Hadoop is no more a simple batch processing platform in all sort of data science use cases. It has further become a multiple purpose engine for doing ad hoc analysis and even has become a tool for daily workload operational reporting.
In the year 2017, organizations may further start responding positively to these hybrid systems also by adopting architectural designs which are use case-specific. A host of factors will be playing vital as the user personas, volumes, questions, access frequency, data speed, and aggravation level before committing to any data strategy.
These new-age architecture references will be more personalized needs driven. Enterprises may be capable of combining the top self-service type of data preparation tools, Core Hadoop, and various end-user analytics tools in such a way to perform flexibly as per the needs of the situation.
6. Machine learning and Spark may further enhance big data
Apache Spark was once an essential component of Hadoop ecosystem, which has now become a choice-based big-data platform for enterprises. In a recent survey conducted on data architects, about 70% of the participants including the BI analysts and IT managers favored Spark over the latest MapReduce, which is primarily batch-oriented and will not lend itself to the interactive apps to serve the real-time streaming purpose.
Big data enthusiasts can also see that a variety of Big Data resources are making real significant discoveries to take the future into new and innovative directions. We will discuss such shifts, inventions, and practical application of big data advancements in the forthcoming articles.