Image optimization techniques are implemented for the purpose of reducing the size of an image file. The goal is to help browsers quickly load images, and ensure user experience remains positive. Another advantage of image optimization is reducing network bandwidth consumption, which can help speed up file transfers and save on bandwidth costs.
The most common way to optimize images is image compression. However, to significantly reduce file size, many are forced to use lossy compression, which often degrades the quality of the image. Machine learning (ML) and artificial intelligence (AI) algorithms are currently being developed and trained for the purpose of improving image optimization and image quality.
Image Optimization Challenges
There are a few key challenges facing image optimization methods and techniques:
- Image format limitations ”image optimization algorithms operate on existing image files. Some image formats have lossy compression, which removes some of the information from the original image to reduce file size. This creates a need for upsampling , ways to increase image resolution and add the missing data back to the image.
- Resizing ”a highly effective way to optimize images is to resize them to the actual size needed by the web page or application. There may be several sizes for the same image, for different devices and screen sizes. However, this can be challenging for large websites with thousands or millions of images. Each image must be resized and cropped to the appropriate size, retaining its essential elements. This process can be automated using machine learning techniques.
- Balancing quality and resolution ”each image format has parameters you can tune to control image resolution, image quality and level of compression. To optimize an image, you need to tune an image to the lowest possible quality and resolution that will still appear sharp to the human eye. This type of tuning is difficult to achieve manually, and is another field of research for machine learning algorithms.
- Image metadata ”some metadata is created by cameras and graphics applications, while other metadata is applied by users. Metadata that is not necessary for the user or application increases image file size, and should be removed. At the same time, some metadata provides information about the original image and can be used in analysis and optimization.
5 Image Optimization Techniques Leveraging Machine Learning
Here are five ways machine learning and deep learning algorithms are used to improve image optimization, in ways previously not thought possible.
Image Compression and Resolution
A new algorithm developed by Google, known as Rapid and Accurate Image Super Resolution (RAISR), uses deep learning and conventional sampling techniques to capture low-resolution images and produce high-resolution versions. It is very innovative, as it was previously not possible to upgrade the resolution of existing low-resolution images. RAISR works by identifying edge features in the image and training on them to generate realistic image enhancements.
Image Manipulation and Generation
This technique uses adversarial neural networks, where one network teaches the other how to generate realistic data representations. This can achieve interesting effects such as turning night into day, or changing the weather in an existing image. The adversarial networks’ ability to learn from each other and generate new, realistic images is likely to be used in a variety of applications, from healthcare to film production.
This technique can be used to increase image resolution. For example, researchers have trained the Super-Resolution Generative Adversarial Network (SRGAN) to predict and populate missing data in low-resolution images, creating photo-realistic high-resolution output.
Image Enhancement in Live Video Streams
Google has created a neural network that can capture frames from high-definition video, run them in real time on mobile devices, and apply image enhancements on the fly. The main innovation is the speed at which these transformations are performed, especially considering the limited hardware available on mobile devices. This new method allows users to modify and apply filters to videos in real time.
Image Compression with Deep Learning
Based on a 2017 paper by Jiang, Tao, et al., this technique uses two CNNs to create an end-to-end compression framework. The first is a ComCNN (Compact Convolutional Neural Network) and the second CNN is a RecCNN (Reconstruction Convolutional Neural Network):
- ComCNN learns the best compression representation from the input image and uses an image codec such as JPEG to decode it.
- RecCNN reconstructs the decoded image in high quality.
This neural architecture can be used for image compression, image denoising, resampling, image restoration, and image completion. It is compatible with existing image codecs such as JPEG, JPEG2000 and BPG.
Upsampling with Deep Learning
In their 2018 paper, Paul Andrei Bricman and Radu Tudor Lonescu propose using the CocoNet neural network for mapping 2D pixel coordinates to corresponding RGB color values. Here is a simplified overview of how the process works:
- Training ”the neural network learns how to encode an input image in the neural layers, using a continuous function.
- Testing ”the neural network is given a 2D pixel coordinate, and then outputs the approximate RGB values of a corresponding pixel.
Once the neural network learns the process, it can perform tasks like image encoding, image compression, image denoising, image resampling and image completion.
Conclusion
In this article I reviewed the field of image optimization, key challenges and how they can be addressed by AI/ML techniques. I also briefly reviewed five approaches to optimizing images using deep learning algorithms:
- Improving image compression and resolution using Google RAISR algorithm
- Performing image manipulation and feature generation using the Google SRGAN adversarial framework
- Enhancing images in live video streams using a high performance neural technique
- Compressing images using deep learning with a dual convolutional structure ”ComCNN and RecCNN
- Upsampling images using deep learning, based on the CocoNet neural network