Posted in

How and where to use Embedding in Machine Learning?

Howdy folks!

Does handling voluminous data to train your machine learning model give you nightmares? Do you know there is a way to find out similar song? Do you know how friendship recommendation in social networks works? I bet; this blog is worth a read then.

Introduction

The world of ML/AI has so evolved that it has started serving the sector of healthcare, security, life sciences and other domains. Eradicating time and manpower, AI business solutions have become self-sufficient. This technology has contributed the various verticals of analytics such as data, text, image and video analytics making itself eligible to work with doctors and other healthcare professionals.

Building deep learning models won’t be a challenge anymore if the right input (data) with the capability of handling huge volumes is there. Sometimes, data is such big that it becomes difficult to train or predict ultimately. Building machine learning models give pain when dealing with huge data to train. Hence, embedding come into picture. Also, embedding again come into picture when you intend to conclude the similarity between two items. Before coming over your curiosity about it, let’s glance at following topics we will go through.

Milestones

  • What is embedding?
  • Usage of embedding in ML
  • How to generate embedding?
  • Applications of Embedding

What is Embedding?

Embedding as a literal meaning speaks of an extract (part) of something. In the context of machine learning, embedding is a low-dimensional data converted from high-dimensional data in the form of a vector in such a way that those are semantically close to each other. Deep Neural Network models can be trained (learned) to generate embedding and same pre-trained model can be reused to generate another embedding for a different set of data. This made machine learning easier to tackle with large dimensional inputs.

Usage of Embedding in ML

In machine learning, embedding can be useful in several of its contexts. This is proved to be very useful in a recommendation system affiliated with a collaborative filtering mechanism. An objective of item similarity use cases is what helps in such systems. Another objective is to keep data simple for training and predicting. Performance of ML model significantly got improved after usage of embedding. The only shortcoming is that embedding makes the model less interpretable. Let’s suppose, 20 categories got reduced to 3-D categories.

Well, in machine learning embedding are to achieve an objective for either dimensionality reduction or item similarity check which will be discussed later.

How to Generate Embedding?

Embedding is generated in such a way that the input data points which are semantically same to each other are placed close. Think in a way like how knn algorithm works. Similar properties data points or items are placed close to each other.

There are various techniques of generating embedding in a deep neural net which entirely depends on your objective.

Objectives are as follows.

  • Similarity Check
  • Image search and retrieval
  • Recommendation system
  • Word2Vec for text
  • Similar songs
  • Reduce high-dimensional input data
  • Categorical Variables in huge number can be compressed instead of using one-hot encoding. For an example, you can refer to this site.
  • Eradicate sparsity; hen most of the data points are zeros, recommended converting those into meaningful lower dimension datapoints
  • Multimodal translation
  • Image captioning

Now, let’s look over the ways to generate embedding. We will not go through all embedding types, instead let’s just focus on image and text embedding.

Image Embedding

1. Using encoder of Autoencoder

Train images using Autoencoder network which is a complete set of encoder and decoder network, a convolution neural nets (CNN). Encoder network is generally decreasing in size while the decoder is the opposite. The output of the encoder is called an intermediate output or latent space representation. This is a low-dimensional representation (compressed form) of input image data which is passed onto decoder.

Source.

This intermediate layer (compressed data) is a trained embedding for the input images and the number of neurons here is the size of embedded vector. Using this encoder network, you can predict (generate) embedding for the new image. Let’s glimpse over the same.

Step1: Import required libraries

Step2: Just to generate embedding, there is no need to have a decoder network. Instead, let’s just build an encoder which we will compile.

Step3: Keras model above is initialized using Model class. Also, encoder network’s architecture is shown below.

Step4: Now, the output of encoder which is a latent-representation of input images or just say embedding is an intermediate layer. This will be useful in generating embedding for other images.

2. Using Pre-trained Model

Train an image classification model on a large dataset. Remove the last layer of sigmoid/softmax activation function. Pass an image or set of images to this last layer removed model to generate single-vector embedding. Later, this feature vector can be used for similarity checks or other machine learning use-cases.

Let’s see its implementation.

Step1: Import required libraries

Step2: Build a CNN Keras model for image classification using 2 hidden layers having relu’ activation each and sigmoid’ function in the last layer.

Step3: Compile the structured model above

Step4: Train the compiled model

Step5: Remove the last layer and prepare a new model without that last layer (flattened)

Step6: Generate embedding for an unseen test image from the trained Keras model

Text Embedding

Google invented Word2vec algorithm for training word embedding. This algorithm maps words in a way closer to each other, sharing similar semantic meaning.

For instance, a book titled Victory in hand is similar to Preparation of Win although they don’t share any word. But since their semantic encoding is very similar thus, algorithm detects titles as similar to each other.

Other Embedding

Others like Doc2vec to create embedding from a document, Gene2vec to create embedding from genomes, or even Graph2vec to generate embedding from complex data structures like graphs and many others.

Applications of Embedding

  • Recommendation Systems
  • Automatic image captioning
  • Linkage-intensive domain: Drug Design, Protein-to-protein interaction
  • Identify similar items such as songs, movies, products, or genomes

Conclusion

Embedding in ML/AI, particularly in deep neural nets, are so much useful in building a model. As described above, immense applications can be built using embedding eradicating various challenges of building models such as treating huge amounts of images or any other format of data analytics services. It has spread its wing among various horizons in healthcare and commercial analytics as well. The blog has deeply illustrated how embedding can be generated using various ways.

I am a full-time guys and a part-time blogger. Daniel Jacob is a globally writer for a big data, artificial intelligence, machine learning, data analytics, python and other emergency technologies. He holds a bachelor of Technology in New York Institute Technology.

Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.