Posted in

Machine Learning: Bridging the Gaps in IT Data Silos

In todays complex business world where many organizations operate in silos, data is plentiful and its challenging to get a big-picture view of the entire IT landscape how can enterprises better manage, analyze and interpret tremendous amounts of data?

The next big thing in ITOA machine learning is providing a viable solution. Machine learning studies how to design algorithms that can learn by observing data, discovering new insights in data, developing systems that can automatically adapt and customize themselves, and designing systems where its too complicated and costly to implement all possible circumstances, such as search engines and self-driving cars.

Theres been a significant increase in machine learning applications in ITOA due, in large part, to the ongoing growth of machine learning theory, algorithms, and computational resources on demand. Many organizations are finding that machine learning allows them to better analyze large amounts of data, gain valuable insights, reduce incident investigation time, determine which alerts are correlated and what causes event storms and even prevent incidents from happening in the first place.

In fact, machine learning can help cure a variety of IT pains with the following technologies:

Clustering

Imagine a large, global corporation, with tens of thousands of servers across their enterprise.  Their IT systems produce a half-million events per hour, generating 50,000 or more help desk tickets per year. In this scenario, they have more than 2,000 level-2 escalations, which calculates to around 50 escalations per day.  How can their IT department manage this enormous volume?

If IT looks at each alert separately, they cant see a full picture about whats happening.  Investigating each ticket individually takes significant time, effort and expertise, and doesnt always immediately identify root cause.

Machine learning can cluster similar items together, automatically identifying meaningful relationships through algorithms. Similarity could be defined as distance in time, host, service etc. Clustering establishes high-level situation awareness by removing redundant, low-quality alerts by clustering alerts into meaningful groups.

Anomaly Detection

Many enterprises already monitor various performance indicators, including server workload, transaction execution timings, application performance, and end-user response time. Usually, they need to be aware of two aspects:

  1. Monitoring potentially bad situations, which are usually specified as policies identifying known problems, e.g., threshold for low remaining disk space alert; and

  2. Monitoring good situations and when they stop happening.

Its important to identify unknown problems, such as deviation from steady/stable state, drop in desired behavior of the system, sudden decrease in performance, etc. Typical approaches rely on dynamic thresholds based on standard deviation calculations, which aim to address this, but in practice such models are too simplistic, causing too many false alerts. For example, without machine learning, occasional spikes in a KPI are reported as alerts.  In reality, such spikes may instead indicate part of the system normal behavior, caused by a daily-run automatic script, for example. This is an expected normal behavior, not something IT should be alerted about each time.

Machine learning can prevent this problem by identifying normal system behavior and reporting any anomalies deviating from it, by constructing behavior signatures and applying an anomaly detection algorithm atop of it. Such algorithm first observes how the system normally behaves and then starts reporting significant deviations from it. Moreover, the algorithm is able to continuously adapt its behavior signature library, thus learning how behavior changes over time.

Causal Reasoning

Understanding how two time series are correlated doesnt imply which one caused the other to spike.  Such analysis does not imply causation, and IT needs a better understanding of the cause-effect relationship between data sources.  Its critically important to understand which data sources contain triggers that will impact the environment, the actual results of the triggers, and how the environment responds to the changes.

Machine learning can establish basic relationships between collected data sources, and correlate events, tickets, alerts, and changes using cause-effect relationships such as linking a change request to the actual changes in the environment, linking an APM alert to a specific environment, and linking a log error to a particular web service, etc. When dealing with various levels of unstructured data, the linking process (or correlation) isnt that obvious, so machine learning can establish relationships between different data sources, determine how to link them to environments, and when it makes sense to.

Frequent Pattern Mining, Classification and Forecasting

Often, IT is reactive to incidents that have already happened, rather than proactively preventing the incidents from occurring in the first place. Incidents typically start with a change in IT system, which causes particular components/services to fail, triggering an alert. At this point, the Tier1/Tier2 team reacts to the incident, starts an investigation, proposes a fix and resolves the incident.

Now, preventive analytics can detect problems early, stopping them from turning into incidents. Machine learning drives preventive analytics with two approaches: frequent pattern mining and forecasting.

Frequent pattern mining automatically crawls data to identify items frequently appearing together. The same algorithm powers retail basket analysis, identifying what items are frequently bought together, whats the next product that customers are likely to purchase, and how to bundle products/services together to maximize revenue. The same way can frequent pattern mining identify changes and alerts frequently appearing together, which can be used to estimate the effects even before a change in deployed.

Next, forecasting can be used in combination with frequent pattern mining to estimate the magnitude, severity, and timing of potential issues. By looking into the past behavior of the system, the models can indicate how systems react to various changes in workload, configuration or infrastructure.

Conclusion

IT operations analytics has a large opportunity to leverage Machine Learning algorithms to resolve common problems, analyze large piles of data, view information across silos, and prevent incidents from occurring. This, in turn, can help IT run far more efficiently, effectively and innovatively.

Bostjan Kaluza is the Chief Data Scientist at Evolven. He's also a researcher interested in artificial intelligence and intelligent systems, machine learning, predictive analytics and anomaly detection. Prior to Evolven, Boštjan served as a senior researcher in the Department of Intelligent Systems at the Jozef Stefan Institute, the leading Slovenian scientific research institution and led research projects involving pattern and anomaly detection, machine learning and predictive analytics.

Focusing on the detection of suspicious behavior and data analysis, Boštjan has published numerous articles in professional journals and delivered conference papers. In 2013, Boštjan published his first book on data science, Instant Weka How-to, exploring how to leverage machine learning using Weka. Boštjan recently published his second book, Practical Machine Learning in Java, for learning how to use Java's machine learning libraries to gain insight from data. Boštjan is also the author and contributor to a number of patents in the areas of anomaly detection and pattern recognition.

Boštjan earned his PhD at Jožef Stefan International Postgraduate School in Ljubljana, Slovenia, rigorously defending a doctoral dissertation entitled Detection of Anomalous and Suspicious Behavior Patterns.

Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.