Introduction
Data Structure is the way of organizing the data to retrieve it with minimum cost and utilization of resources.
On the flip side, Machine Learning is a field of computer science that focuses on the use of data and algorithms to intimate the way of learning.
Machine Learning overall consists of approaches and techniques which are entirely built on statistics, probability and optimization.
The first two building blocks are related to mathematics and the third one is related to Data Structures and Algorithms. Ultimately Machine learning is a field modeled to play with data and generate something significant.
What is the use of Data Structures in Machine Learning?
1. The link Between Data Structures and Machine Learning
Basically, the essence of Data Structures is how we store data and retrieve the data. Programming language is a medium to represent those structures in a human-readable way.
Now assume that there is a problem that we want to solve using machine learning.
Then as a Machine Learning professional, you need to be aware of which model is fastest and eats up minute space while precisely solving the problem.
This model often consists of steps that are using multiple data structures to achieve the above-mentioned objectives.
So a professional having a good hold on Data Structures can answer the following question that he/she has to face in daily work.
- How much time will the model(solution) take to complete the process?
- How many space resources are utilized while doing the process?
- Which model is better while considering the trade-off between time, space and business requirement
If a professional is working in production then a terrific grasp of data structure, algorithms and computer architecture is necessary to drive business solutions.
2. Real-time Predictions in Machine Learning
Assume we have a problem of object detection which we want to solve using machine learning. To solve this we have a model where we are getting 10 frames per second as input and our algorithms in the model will accumulate those frames to generate the desired output.
Our model has a requirement of a minimum of 10 frames per second which we can call real-time input. In the worst case, If the input rate goes beyond 10 frames then input can be classified as obsolete and model prediction can be seen as laggy and it wouldn’t be able to give the desired output.
So if a practitioner has knowledge of Data Structure and Algorithms then he/she can easily modify algorithms with the use of proper data structures to improve performance up to the mark. Which will further result in object prediction in real-time.
3. Link Prediction Machine Learning Algorithm
We will take the example of social media, Suppose we want to update you with suggestions of who can be your next connection.
This problem can be easily modeled as a graph data structure where there are 2 entities and we want to figure out if there is any link between them.
So primarily you need to model a person as a node and the connection between two persons as an edge then you need to create a graph of them or precompute it as per minimum utilization of resources.
Then using the BFS/DFS graph traversing algorithm we need to check if we can visit the second node after starting from the first node. This graph data structure has a huge influence in the machine learning field whenever there is a problem with entities having relations between them.
4. Hashing in Machine Learning
Now suppose we have an enormous data set that may consist of duplicates. On top of that, we are getting records as a stream.
In this case, normally professionals will think that each input record will graze over all available records and if there is any record that is the same as the input they discard the input.
But if we consider the time it takes for each input is linear because for each input we are visiting all the available records.
Here Hashing comes into the picture which will reduce this searching time from linear to constant. So whenever a record comes we will convert the record into a hash value then we
will confirm if anything is there at that hash value if yes then we can say this is a duplicate else we will add it. Primarily use of a hashmap or set will reduce the time required for searching drastically to asymptotically constant time.
5. K-way Merge in Machine Learning
Now think of a use case where we have to design the machine learning algorithm where we are getting sorted streams from K multiple IoT devices which act as input. then our model generates a single sorted stream from K streams.
Here, Heap data structure comes to the rescue. In short, Heap is a data structure in a complete binary tree that returns a running minimum or maximum among the stream. Whenever there is an input record we will insert that into minHeap and the second step is to get the minimum from the heap and insert it into the output sorted stream.
6. Machine Learning deployable IoT devices
There are some edge devices that are responsible for properly working for the network. Arduino and Raspberry-pi are some of the widely used IoT devices in the industry. Practically speaking as of now a lot of machine learning algorithms are really heavy to deploy on those devices. Due to these reasons, various top tech companies in the industry are working towards the objective of reducing the time and space complexity of machine learning algorithms. Without knowledge of data structures and algorithms professionals can’t write the optimized code which can be deployable on the edge devices.
7. Unavailability of libraries to solve the problem
While working in computer science as a professional you will encounter problems that can’t be solved using the existing libraries. On the other hand, there is a possibility that you only need one function of the library in the entire application lifecycle so it will result in additional unwanted weightage of the remaining library because the library will be loaded completely.
In the first case, assume that we have data in the form of a tree and we want to visit it level by level. Suppose there is one level after certain steps which is having more nodes than deque can accumulate at an instant. This can result in the breaking of the algorithm. In such a case as a professional, you should have the knowledge to implement deque which will accumulate max nodes in the tree at any level.
In the second case, Suppose you want to deploy the code on the IoT device which needs only one function from the NumPy library. Then there is no point in loading the whole library just for only one use case when we have a space shortfall on an IoT device. So here also
professionals should be aware of data structures and algorithms to implement a single custom function and import only what is necessary.
Summary:
1. As a professional having a good hold on Data Structure and Algorithms is a prerequisite in a machine learning career.
2. Multiple data structures can be used to design machine learning models which can determine the internal details of algorithms.
3. Right choice of Data Structures can optimize the time and space complexity of any machine learning algorithm. for eg. using graphs for object prediction algorithms.