Since it was launched, Kubernetes technology has transformed the way software developers build and deploy applications. The widespread adoption of this technology intrigued data scientists, who saw that Kubernetes also provides features to optimize and support the data science workflow.
In this post, I’ll explain what is Kubernetes and how it can help support data science activities. I’ll also provide some examples of use cases to demonstrate the various benefits that made data scientists fall in love with Kuberenetes technology.
What Is Kubernetes?
Kubernetes is an open-source platform designed to manage containers and clusters in a single interface. You can deploy containers to clusters across all types of environments, including clouds, virtual machines and physical machines, creating a network of mini virtual machines. In Kubernetes, one or more containers are placed in a pod, which is the smallest possible unit that can be deployed.
The platform lets you scale applications according to your workload. It is extensible, in that it allows application components to move across systems.
Kubernetes main features include:
-
Automating manual processes ”manages container hosting and deployment.
-
Self-monitoring ”checks the health of containers and nodes
-
Horizontal scaling ”you can scale out the container applications to accommodate fluctuating workloads.
-
Flexibility ”runs everywhere, on-premises, hybrid or cloud infrastructure, allowing you to move workloads across environments.
The rapid adoption of Kubernetes generated a large active community of users who constantly release features. Some of those features that are useful for data science include:
-
Declarative deployments ”you can establish your production environment in a staging environment.
-
Ubiquitous monitoring ”you can monitor the performance and metrics about any component of a system.
-
Continuous integration ”you can shift from a test suite to a code running in production.
-
Flexible service routing ”you can roll out updates and scale out services.
Kubernetes is more than a container orchestration service, running all categories of workloads. There are many Kubernetes managed solutions that help organizations make the most of implementing Kubernetes. You can learn more about all services Kubernetes provides for enterprises in this guide.
Why It Is A Good Choice For Data Science
Data scientists share similar challenges with software engineers, such as repetitive tasks or experiments. They need to monitor and track metrics in production, manage access and credentials, and scale out with ease. Data scientists can take advantage of Kubernetes capabilities, such as:
-
Continuous batch jobs ”this involves coordinated stages such as process data and test, train and deploy models, usually found in machine learning pipelines.
-
Microservice architectures ”this provides for a simplified application structure based on modularity, which makes it easier to modify and protect your software components.
-
Declarative configurations ”these make it easier to create models across platforms by illustrating the connections between services.
Kubernetes aims to simplify container management by allowing developers to build customized workflows. This feature can prove very useful for data scientists needing to construct dedicated workflows for each experiment.
Kubernetes also provides a framework to build high-level tools such as the Binder service. Binder builds a container image by using a Git repository from Jupyter notebooks. Next, it uses an exposed route to launch the image in a cluster of Kubernetes, enabling to access it via the internet.
Machine learning engineers can also benefit greatly from Kubernetes features.
An example is the Kubeflow project that allows engineers to run frameworks such as JupyterHub, TensorFlow, PyTorch and Seldon under Kubernetes. This allows a machine learning engineer to develop a real portable workload.
Another interesting feature is the integration with Spark, where the software creates a Spark driver in a Kubernetes pod. The driver then creates executors that run in Kubernetes pods, connecting to them and executing applications. You can find more information about how to run Spark in Kubernetes here.
Use Cases
Data science teams can use Kubernetes for numerous applications, such as deploying models for online inference. Kubernetes helps simplify the application scaling to support an increased load by exposing the models so others can use them. Here is how to do it:
-
Create a deployment ”specify which containers you want to run and how many replicas of the application you need to create. You can find the instructions here.
-
Expose the deployment ”use a service to define the rules for exposing the pods, both to each other and to the internet. You can find how to do it here.
Kubernetes will then load-balance the traffic across the replicas. You can configure it to autoscale the resources to meet sudden increases in the workload.
Another example of a use case would be a lab team that needs to analyze data from research and development. Using Spark native integration with Kubernetes, they can access a self-service big data analytics platform.
The use of containers orchestration is especially useful for life sciences research teams. Containers allow for the replication of scientific testing, so scientists can use the exact same software to replicate test results on different environments and devices.
The Bottom Line: What Data Scientists Can Get Out of Kubernetes
The widespread adoption of Kubernetes across the tech industry, notably among DevOps teams, is a testament to its usefulness for optimizing data science and development processes. Kubernetes’ scalability helps data engineers adjust their machine learning workflows with ease, while its flexibility allows them to deploy their workloads across different environments.
At the end of the day, Kubernetes can support data science workflows, optimizing it the same way it does with the software development lifecycle. Therefore, with the staggering amount of big data that engineers need to manage, it is a sensible decision to adopt a Kubernetes solution for data science.