Audio Analytics deals with analyzing and interpreting audio signals captured by devices. Audio data is one of the unstructured data that is in abundance, however, is hugely underexploited. Analyses of this data present an opportunity for various applications ranging from media content analysis, to medical aids to increase CSAT in Customer Service Calls.
In this article, we will explore some of the fundamentals concepts that you need to know in audio data analytics.
Speech vs Voice Analytics
One of the most popular applications of audio analytics is in speech and voice. These are common in call centers where a large amount of audio data is generated and analytics. This represented Call Analytics and involves both Speech and voice. However, note that speech and voice analytics are different tools.
Speech Analytics involves analyzing what is spoken works by analyzing phonetics and involves converting speech to text. The main application of speech analytics is to understand conversations. In contrast, voice analytics deals with how it was spoken “ like syllable emphasis, tone, pitch, tempo, rhythm etc and provides information about the sentiment of the speaker. Both acoustic modeling and language modeling make up most of today’s speech recognition algorithms. In this post, we will focus more on the Voice aspect of audio analytics.
The Central Idea
Sampling is the reduction of a continuous audio signal into a series of discrete values. The sampling frequency/rate is the number of samples taken over some fixed amount of time. The most common sample rate you’ll see is 44.1 kHz, or 44,100 samples per second.
The Sound wave can be represented in an array after the sampling into an array [-4, 0, 4,7,6,3,0,-1,1,3,4,3,-1,-5,-8,-7,-4,-1,0,-2,-4,-5,-4]. However, this is only one way of representing audio data.
Every audio contains a set of many features that represent various characteristics of an audio that we are trying to solve. The frequency-domain representation allows you to observe several characteristics of an audio that are either not visible at all when you look at the audio in the time domain. E.g. if you want to look at the amplitude of frequency or if you need to look at cyclic behavior.
Spectral Features (frequency-based features) are obtained by converting the time-based signal into the frequency domain using the Fourier Transform, like fundamental frequency, frequency components, spectral density, spectral flux, spectral centroid etc.
Finally, let’s look at the most popular one – Mel Frequency Cepstral Coefficient (MFCC).
Mel Frequency Cepstral Coefficient (MFCC)
Mel Frequency Cepstral Coefficient (MFCC) is popularly used in automatic speech recognition. Before MFCC, Linear Prediction Coefficients (LPCs) and Linear Prediction Cepstral Coefficients (LPCCs) were popularly used for audio analysis. MFCCs are coefficients that collectively make up a mel-frequency cepstrum (MFC). MFCCs represents the human voice with the help of a small set of features that describe the overall shape of a spectral envelope.
Let’s look at it from a layman’s understanding. For that, you need to understand cepstrum – which is the information of the rate of change in spectral bands (a sequence of numbers that characterize a frame of speech).
The first step in MFCC is to derive the obtain the Fourier transform of the audio signal. In simple words, obtain what makes up an audio signal. Then map the powers of the spectrum obtained above onto the mel scale – a scale that relates the perceived frequency of a tone to the actual measured frequency. It scales the frequency to match more closely what the human ear can hear (humans are better at identifying small changes in a speech at lower frequencies). The result obtained is the Mel spectrum which is obtained after passing the Fourier transformed signal through the Mel Filter Bank.
The Mel Spectrum is represented in a log scale (humans don’t hear loudness on a linear scale). The Discrete cosine transform (DCT) is applied to the transformed Mel frequency coefficients that produce a set of cepstral coefficients. The DCT decorrelates the Mel filter-bank energies (since they are all overlapping) which means diagonal covariance matrices can be used to model the features. The amplitudes of the resulting spectrum are MFCCs.
This is just a high-level explanation of MFCC, but the process involves several processes. You might also want to learn about variations in the process like the addition of dynamic features like Deltas and Delta-Deltas coefficients.
Hidden Markov models (HMMs) are the commonly used classification models in speech recognition. Alternatives are being researched like Connectionist temporal classification (CTC) systems (comprised of recurrent neural networks and a CTC layer) and attention-based models. You might also want to check out Fréchet Audio Distance (FAD) introduced by Google AI in 2019 to measure the quality of audio from deep-learning networks.
So, before you begin with your first Audio analytics project, it’s a must to understand the fundamentals of the components that make up an audio signal. Most automatic speech recognition (ASR), computer speech recognition, or speech to text (STT) rely on these, and hence it’s important to understand the concept to interpret the output in a meaningful way.