The financial services industry is fundamentally an information processing industry. An investment fund processes information in order to evaluate investments, an insurance company processes information to price their insurances, and retail bank processes information in order to decide which products to offer to customers. It is, therefore, not a surprise that the financial industry was an early adopter of computers.
Machine learning skills are essential for anybody working in financial data analysis. You can use machine learning to automate manual tasks, identify and address systemic bias, and find new insights and patterns hidden in available data.
Machine Learning for Finance shows you how to build machine learning models for use in financial services organizations.
Let’s see what Jannes has to say about using machine learning and AI in the field of Finance:
1. How is machine learning used in finance? How is AI used in finance and banking?
There are many applications that usually revolve around either forecasting or information extraction. Financial forecasting is as old as the industry itself. Analysts try to forecast stock movements or even specific line items to make recommendations to their customers. This is currently done with models that are quite restricted, run on a lot of assumptions, and require a good deal of guesswork from the modeller. Advanced ML methods can replace many of the assumptions and guesses with models learned from data. In a field awash with data, this is very attractive.
Information extraction is newer: next to hard data, e.g. balance sheet data or macroeconomic time series, which has traditionally been used, there is an increasing interest in using soft data, such as news articles, social media posts, or satellite imagery. Extracting useful application from this data usually requires some kind of model. For instance, satellite images of oil tankers are not very useful, but information on how much oil they carry is. So you need a model to extract that information, and machine learning is a good tool for this.
2. How does machine learning facilitate Fraud Detection?
Fraud detection is this constant cat and mouse game between fraudsters and institutions as well as law enforcement trying to stop them. It is really hard to come up with a system of rules that cannot be gamed by the fraudsters. For instance, if you say that a transaction must be less than $10,000, then the fraudster will send $9,999.
A different approach is to look for patterns common to fraudulent transactions. If, for instance, an account usually processes only small payments and suddenly tens of thousands sweep through it in one transaction, it might be worth looking into what is going on. ML can pick up on such patterns through supervised learning. In hindsight, it usually becomes clear when a transaction was fraudulent, and we can train models to detect this.Â
3. How do you use generative models to extract useful financial information?
Generative models such as generative adversarial networks is a relatively new subfield of ML whose aim it is to generate new data that is very similar to existing data. This is especially useful when we lack certain data. For instance, in the fraud detection example, there are (luckily) not that many fraudulent transactions.
Our model would do better if there were more. So we can use a generative model to generate more transactions that are similar to real fraudulent transactions. In the best case, the generative model captures the main underlying features and can generate samples that look very different from the existing samples, but come from the same underlying distribution.
4. What is the black box problem in machine learning? How do you crack open the black box problem in AI?
Modern ML models come with hundreds of millions or even billions of parameters. We are currently unable to understand what exactly the values of all these parameters mean and how they work together. These models do not have an easy way for us to interpret them as, for example, linear regression does. Understanding and interpreting these models is certainly a field of active research. But what we can do here is to measure how these models respond to small differences in the input features.
One method that does this is the LIME algorithm. It changes the features of an input sample and then fits a linear regression model to predict how the model reacts to changes in the sample. This linear model is then very interpretable, so we can use it to make statements like all else being equal, if the applicants’ income was 20% higher, the loan would have been approved. It is worth keeping in mind that such explanations are only estimated explanations.
Usually, if you make 20% more, other things in life change too. So we don’t really know if the model was only discriminating on income or if it was inferring some other features too. But such explanations are a good start to get an idea of what is going on and to give confidence to decision-makers when they rely on ML models.
About the Book
Machine Learning for Finance explores new advances in machine learning and shows how they can be applied across the financial sector, including in insurance, transactions, and lending. It explains the concepts and algorithms behind the main machine learning techniques and provides example Python code for implementing the models yourself.