Abnormal data trends rarely occur on their own. Influencing or related metrics are usually involved. For example, let’s say that a remote data center goes offline and doesn’t come back up. The anomaly in this case isn’t just a power failure, it’s a power failure plus a failure in a backup generator.
Some systems might show you one of these anomalies, leaving you to search for other affected metrics which can take hours, days or even weeks. Correlation, on the other hand, instantly lists related anomalies so you can quickly and painlessly understand what the leading dimension is and which metrics are impacted.
If you don’t use anomaly detection, you won’t understand the cause of your outage until a support crew reaches the site. With anomaly detection, however, you can quickly discover both related anomalies, making it that much easier for you to get back online.
Finding related metrics and anomalies
Behavioral topology learning provides a method for data scientists to understand relationships between millions of metrics at scale. This lets them combine related anomalies into stories, mitigate errors and examine their root causes. When implemented correctly, this system can filter out unrelated metrics from the results for greater accuracy.
There are several methods within the larger category of behavioral topology learning, each with their own advantages and downsides. The challenge is selecting a method that can reliably find related metrics at a large scale.
Correlating simultaneous anomalies
With anomalies, a single problem will usually show itself by displaying multiple abnormalities. For example, a DDoS attack may show itself with both an increase in average latency and in the number of failed connection attempts.
Automatic anomaly detection systems don’t know that two simultaneous anomalies are coming from the same error. What it can do is infer the likelihood of what happens, especially if two simultaneous anomalies recur often in the same metrics. In other words, if two metrics have the same anomaly once, it’s a coincidence. If it happens twice, it’s suspicious. If it happens a third time, the metrics are probably related.
There are many ways for automated detection systems to detect abnormal-based similarities in this way. We find that the Latent Dirichlet Allocation (LDA) algorithm works best. This is because it assumes that a single metric can be part of multiple groups, allowing it to achieve a greater level of accuracy when describing relationships between metrics. One disadvantage of LDA, on the other hand, is that it does not scale as well as other methods. It needs to examine a large amount of historical data in order to function properly.
Names similarity
In data science, having a consistent naming convention can produce patterns of related metrics on its own. For example, the name of one e-commerce metric might be comprised of the name of a web browser, the number of abandoned shopping carts and the location of the customers being monitored: browser=Chrome.Country=US.what=abandoned carts .
If you’re also monitoring a metric called browser=Chrome.Country=Canada.what=abandoned carts, then it’s reasonable to assume that because their names are related, the metrics are related as well.
Normal behavior similarity
Normalcy is subjective. Most metrics look as though they are formed from random spikes, even under normal conditions. In order to discover normalcy at scale, use a machine learning anomaly detection system that identifies patterns out of apparent randomness. When two metrics have the same pattern, they’re probably related.
This method runs into problems when you consider that you can find similar patterns in almost any pair of metrics if you try hard enough. False-positives will abound unless you refine your methodology.
If you want to use a traditional measure such as a Pearson correlation coefficient, your first imperative is to de-trend’ the data. For example, if your revenue is constantly increasing, you need to remove its upward trendline before trying to correlate it. Otherwise, any metric that is also trending up will be falsely correlated with your increasing revenue. You should remove seasonality for the same reason; any two metrics that have the same seasonal patterns will also be falsely correlated.
If you want to do less work and encounter fewer false positives, try a pattern dictionary approach. Almost every time series can be divided into a number of archetypal patterns, such as a sine way, a sawtooth, a square wave and so on. If two metrics assume the same shape at the same time, it’s likely that they’re related.
User input
Like with naming, user input is one of the less scientific methods of understanding which metrics are related. Groups of metrics are essentially related if an authoritative user simply says that they’re related. This fact can be encoded into the machine learning model. In addition, if it makes sense to create a composite metric “ such as the sum of abandoned shopping carts in every country “ then the constituent metrics of that composite metric are likely to be related as well.
Locality Sensitive Hashing (LSH)
Naming groups of metrics or using user input might be highly accurate, but it’s also a tactic that doesn’t scale to millions or even billions of metrics. Meanwhile, algorithm-based methods need either a lot or time or a lot of computing power.
So how do you make the work faster and less computationally-expensive?
One way is to divide and conquer. Take a billion metrics and then divide them into 100 groups of approximately related metrics. That leaves us with 10 groups of 10 million metrics. This is a lot smaller than one billion, which means that it would be computationally easier to compute the similarities between them. The question, of course, is what mechanism do you use to undertake the division?
Locality sensitive hashing (LSH) is a lightweight computation that tags every metric that a company runs and assigns it to a group. Once the metric is assigned, you can run some additional expensive computations on them. Although LSH has the possibility of grouping metrics that shouldn’t be related or ungrouping algorithms that should be related, the system can be turned over time.
Putting it all together
Combining the algorithms and methods above helps provide a more useful picture of your platform. Any single anomaly is not likely to point to just one issue. Correlated anomalies that can be identified across multiple anomalous metrics are truly worth the attention. By using the metrics above, you’ll be able to minimize a storm of anomalies into an issue that’s much easier to solve.