Posted in

The Human Microbiome and its Application to the World of Data Science

How do you feel knowing that there are other living things having a fun-filled day on your skin and inside your guts? In fact, different regions of our body have a distinct number of organisms living on them which differs from one person to another. Examples of these microorganisms are bacteria, archaea, fungi, and viruses. The sum-total of all these microorganisms living on and inside us is the human microbiota. They are super important because they help us fight diseases, absorb nutrients, and even contribute to mental health. Now, the genetic material of all these microbes that live on and in the human body is the human microbiome. These genetic materials are responsible for how the microorganisms in the body behave. 

The human microbiome is also very important to our health because its dysfunction can lead to some autoimmune diseases. In an autoimmune disease condition, the body’s immune system attacks healthy cells. Examples are diabetes, fibromyalgia, rheumatoid arthritis, multiple sclerosis, and muscular dystrophy. When the balance of the microorganisms in the body becomes interrupted due to factors such as dietary changes, new antibiotic medication, accidental chemical consumption, poor hygiene, or high levels of stress, disease-causing organisms build up and invade the normal microorganisms changing their genetic makeup. This disruption causes an abnormal immune response against the substances and tissues in the body. Thus, autoimmune diseases are passed by inheriting the family’s microbiome.

Why is Data Science Applicable to Human Microbiome?

 

Microbiome studies produce lots of data that require high-level computational tools to be analyzed. These enormous data sets present new challenges and opportunities for computational biologists to understand, and even manipulate the human microbiome to improve human health and data science is the best tool to use in analyzing them. To study the human microbiome, animal models are used extensively because they allow control of specific variables that are unachievable in human studies. To test whether the microbiome composition and function correlate with host variables (age, diet, genotype, phenotype, etc.), or experimental exposures, less expensive animal models such as fruit flies, zebrafish, and Caenorhabditis elegans can also be used.

Why Does This Matter to Us?

The human microbiome is important to us because it evolves over time in response to age, nutrition, medication, and even diseases like cancer. Researchers need to track how the expression of these bacterial genes changes over time and how it differs among patients by linking microbiome sequences with patient blood tests, epigenomic data, histological images, and even clinical findings to make sense of the microbiome.

Data Visualization and Statistical Methods for Microbiome Analysis

The approach to analysis in most microbiome studies is to look at differential microbial diversity, functional components (gene or biochemical pathways), or taxa abundance between the comparison groups. Because microbiome datasets are high-dimensional as they can potentially have thousands of taxonomic units common statistical methods cannot be applied. 

Graphical Network and Machine Learning in Microbiome Science

Graphical Network in Microbiome Science

Graphical networks are used to compare microbiome interactions in different states, for example, healthy state vs diseased states or to illustrate which organisms coexist or are mutually exclusive to each other. Normally, networks can be
inferred using pairwise relationships where networks are based on similarity or
correlation coefficients between pairwise variables. Commonly used programs to
infer correlation networks for microbiome data are SparCC (Sparse Correlations
for Compositional data), CCLasso (Correlation inference for Compositional data
through Lasso), and SPEICEASI (Sparse and Compositionally Robust Inference of
Microbial Ecological Networks). Among commonly used programs to infer correlation networks for microbiome data SPEIC-EASI (Sparse and Compositionally Robust Inference of Microbial Ecological Networks) is the most widely used model as it appears to be the most robust estimating interactions. 

Machine Learning Methods in Microbiome Science

Because they can be used on many different types of high-dimensional data and are especially useful for feature selection in multi-feature datasets, machine learning algorithms are emerging as a widely accepted tool in the analyses of microbiome datasets. For example, Random Forest is a machine learning technique for classification and regression that is often applied to identify important taxa and clinical covariates and can differentiate various phenotypes or forecast precise results. Although methods such as CART analyses (Classification and Regression Trees) are less accurate, they are more interpretable and clinically actionable as decision trees allow the investigator to have deep understandings of which variables are important. It is important to note that these models must always be cross-validated either by enrolling separate cohorts for training, testing, and validation or via sample and replacement.

What Does the Future Hold for Human Microbiome Science and Data Science?

 

Although microbiome data is complex given the standards of today’s data-driven world. The pace of research is accelerating as microbiome researchers get funding and build new databases that others can use to answer new questions as with most data science fields. While this happens, work balance in the microbiome space is moving from the confines of the laboratory towards the data science world. 

Oladimeji Ewumi is a freelance health and AI writer who is passionate about helping healthcare, AI and B2B brands communicate with target audiences through effective story telling. 

Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.