Machine learning depends on labelled data to train an algorithm to detect patterns. In the computer vision space, these algorithms are trained to detect objects, recognize activities, or perform segmentation. A good, representative, labelled dataset is of vital importance for training such computer vision algorithms. In this article, we will explain data collection, manual labelling, automated labelling that we perform for our projects in Optisol Data labs.
Data Collection
Data collection is an important aspect of machine learning. Typically, data for computer vision algorithms are collected as video feeds or individual images. When data the algorithm needs to be trained is proprietary, we depend on our clients to give us representative training dataset for us to label and train the models. If the data can be collected from publicly available data sets like Google image search or Kaggle, we will collect them ourselves. But this is easier said than done.
There are a lot of obstacles to collecting good data. There may not be enough cameras or other sensors in the right location to collect the data we need to train. There could be environmental factors like how the factory floor and workstations are set up that may occlude the view of the cameras. In the fields of petrochemicals refining and other places where a fire is a very dangerous hazard, the cameras need to be intrinsically safe. Such cameras are very expensive to purchase. If we are tasked to collect our own data, which is quite often, there aren’t enough representative samples that we could collect. For these and other reasons, customers or ourselves may not be able to acquire a good dataset. In such scenarios, what do we do?
We have started to use 3D modelling and animation to simulate the environment from a few reference images. Such models give us the ability to generate as big a dataset as our model training requires without expensive and time-consuming site visits and coordination with clients who are located far away.
Following are links to our youtube channel where we have uploaded some of the 3D animation videos that we have created for the purpose of data generation for model training.
Data Labelling (Manual)
Once the data is collected the next tasks is to label the data. We use a slew of data labelling software to label data. LabelImg, LabelMe, VGG image annotator and Some of these tools are quite straight forward use. We can draw polygons like squares or rectangles that on the images and the tool will output the labels and their positions in a format that allows us to train a model.
However, there are other labelling requirements like segmentation algorithms that are quite trick and intensive to generate labels. These labels need to match the contour of the object or segments that we are training the model to differentiate. So, tracing the outlines of such objects on a big data set is a labour-intensive operation. Sometimes we can’t avoid it. Other times we try to use automated data labelling to help us out. Here is an example of labelling we have done to detect and segment an electrical utility pole that we did for an overseas client of ours

Data Labelling (Automated)
Data Labelling is an important, labour-intensive process in the model training pipeline. To automate this, we have employed various strategies. The one that works the best is to tag enough data to train the first iteration of the model. Once the model is trained, we use it as an automated tagger. We run more data through this model and let it detect the objects it is trained for. We will review the model output and manually re-tag the incorrect inferences.
These mistakes that the model makes informs us of some of the structural deficiencies in our training pipeline. This helps us to adjust the pipeline accordingly. The curated output from both the model and the manual re-tagging will be fed back as training data to train the next iteration of the model. We will keep repeating process this till we get a model that meets the requirements of the client.
As the model evolves, there will be less and less data that needs manual re-tagging. This also allows us to farm out the tagging process over a period of days and weeks which helps smoothen out the resourcing curve for our projects.