Beauty of machine learning lies in its power to generalize. Generalization is the algorithm‘s ability to learn patterns in a given dataset and predict results. For example, if an algorithm has been trained to predict mortgage price based on a borrower’s credit history and annual salary, that information can be used to predict any new mortgage prices. However, if the same machine learning algorithm is used to predict car insurance price, it will fail. The developer will have to start the process all over again. They will need to acquire a new set of labelled training data, understand the insurance domain, build and train their algorithm and finally deploy it to production.
Supervised learning process described above is computationally and financially expensive to extend into another domain. So, how can a company without enough resources utilize the power of machine learning to build an algorithm and make predictions in a different domain? Luckily, there is Transfer Learning, an algorithm design methodology [1] where an algorithm that has been trained to understand knowledge from one domain can be applied to a different domain.
This is a developing field and Dr. Andrew Ng, former Chief Scientist at Baidu, states that it will be the next driver of machine learning success across industries. This is true as industrial systems change with time and newer technologies are introduced. With newer technology, newer data acquisition method and data types emerge. However, application to analyze newer data type might still have older machine learning models associated with it. The application will then fail to perform optimally causing various prediction issues. Hence, application with transfer learning algorithm designed into it would be ideal for evolving systems.
Currently, transfer learning has been extensively used in image recognition and some natural language processing application. In image recognition, deep-nets are trained on available ImageNet dataset and then deployed to production. This way, an algorithm trained to recognize certain labelled images can be used to recognize other unseen images of a different variety. A simple overview of a process to architect transfer learning system is described below.
When designing such system, we need to define a source task and a target task. In source task, deep learning algorithm is trained on readily available labelled data. In target task, that algorithm is applied to a different but related task. For example, if a convolutional neural network (CNN) is trained on dataset to recognize animals, this becomes a source task. In this case, the entire neuron layers are trained with a variety of labelled animal images. This trains the algorithm to recognize any new animal images. If we were to extend the same model to recognize a vehicle but don’t have enough labelled vehicle images, retraining the entire CNN layers will mess up the learning. So, we add a few adaptation neuron layers towards the end of our previous CNN and train this new layer with available labelled vehicle images for the target task [2]. This fine tuning ensures we are not overfitting the entire network to recognize only few vehicle images, thereby, retaining the algorithm’s generalization power.
That is just one example of transfer learning and a high-level overview of designing a network. There are other potential applications as well. Researcher Sebastian Ruder lists few in his article where transfer learning can be used in real world problems. For example, training self-driving cars by using data from self-driving simulations [3] is one. Given the flexibility of using a trained model for variety of similar task, it is recommended that businesses invest their research and development efforts in transfer learning.
References: [1] Hulstaert, Lars. (2018) Transfer Learning: Leverage Insights from Big Data [Online]. Available: https://www.datacamp.com/community/tutorials/transfer-learning [2] Oquab M., Laptev I., Sivic J. and Bottou L. (2014) Learning and Transferring Mid-Level Image Representations using Convolutional Neural Networks, 2014 IEEE Conference on Computer Vision and Pattern Recognition. [3] Ruder, Sebastian. (2017) Transfer Learning - Machine Learning's Next Frontier [Online]. Available: https://ruder.io/transfer-learning/