Posted in

Everything You Must Know about Hugging Face’s BigScience Large Open-science Open-access Multilingual Language Model (BLOOM)

Under Hugging Face’s BigScience initiative the world’s largest open multilingual language model BLOOM has been officially launched this July.

The initial version 1.3 has as many as 176 billion parameters to generate text in 46 natural languages and 13 programming languages. BLOOM is the first language model with over 100B parameters ever created for languages such as Spanish, French, and Arabic. BLOOM model is a result of over a year of hard work that involved more than 1000 researchers from 70 plus countries and over 250 institutions leading to the final run of 117 days of training (March 11 – July 6) wherein the BLOOM model was trained on the Jean Zay supercomputer in the south of Paris, France before it was officially made available for researchers.

Hugging Face is an open-source AI community focused on building futuristic solutions with specialties such as machine learning, natural language processing, and deep learning.

Two prominent French agencies CNRS and GENCI provided a compute grant worth 3M million for the research and training of the BLOOM model.

BLOOM is designed to enable researchers to seamlessly access a variety of documents and semi-structured data, discover and select the best results from many documents, capture the structure of these documents into text corpora, and analyze these documents in order to improve machine translation. BLOOM can also be instructed to perform text tasks it hasn’t been explicitly trained for by casting them as text generation tasks.

This blog post covers everything one must know about Hugging Face’s BigScience Large Open-science Open-access Multilingual Language Model (BLOOM).

Hugging Face’s Vision Behind Creating BLOOM

In AI, large language models (LLMs) are powerful general models that excel at a wide range of language tasks. They can take on new problems without modifications to the model architecture, but they’re incredibly difficult to study or even use due to the exclusive access of only a few industrial labs.

BLOOM is a multilingual language model (LLM) designed to tackle challenging applications. It’s the first major open source project to authoritatively evaluate and rank human-level performance in a range of languages, with ground-breaking tools for evaluating them.

Hugging Face’s BigScience project created and made the world’s largest open research project available to the public in the form of BLOOM. It is a complete, unbiased, and multi-lingual model which can be used by all you want and by many different individuals, companies, and institutions.

At a time when more successful startups are coming from large AI research collaborations, Hugging Face understood the importance of developing the field of open science and transforming scientific publications from a record of compiled projects to a record of science itself, and successfully took a first major step towards it as BLOOM went live. As this model is open-source, all information including technical details, model architecture, objective, compute infrastructure, and ideal and intended ways of use are available on the official website of BLOOM.

A new addition to the HF family, BLOOM (and its successor) represents the deepest dive into language modeling and the largest publicly-available language model repository. In this way, Hugging Face is opening up our research to everyone. BLOOM can use it directly to build high-quality models that are used by industry and individuals across a diverse range of domains. The Hugging Face team believes that this will be the definitive high-performance support vector machine (SVM) library for deep learning.

The power of BLOOM is being shared with the world and the Hugging Face Team is elated about it. They see it as the first step in a long, exciting journey-one that will continuously improve and make more data available for more people. They have set out to develop and open source a novel approach that could solve the problem of language model training speed, while at the same time keeping it safe from cheating. Their goal is to provide all users of large language models with the ability to easily build custom approaches and tools, whether commercial or open source.

How Bloom Will be a Valuable Asset for the Researchers?

Bloom is an autoregressive language model. It has been trained to continue text from a prompt on large amounts of text data. It leverages industrial-scale computational resources to provide coherent text as output in 46 languages and 13 programming languages. The texts generated by BLOOM are hardly distinguishable from those written by humans. Moreover, BLOOM can also perform tasks related to the text that it hasn’t been exclusively trained for. All one needs to do is cast these tasks as text generation tasks.

How to Deploy BLOOM?

One can use HuggingFace’s ecosystem to deploy the model and use it further. One needs to have transformers and accelerate installed. The model can be downloaded as follows:

BLOOM is now available for everyone to use and study. One can get started by downloading, running, and studying BLOOM – the Hugging Face team is excited to see what innovators come up with the aid of BLOOM. The model is still being continuously improved and optimized by the Hugging Face team. They are in the midst of finalizing an interface API for large-scale use. For quick-test, prototyping and lower-scale use one can deploy a primitive version on the HF hub.

The Intended Use of BLOOM

The model will be used for public research on large language models (LLMs). This includes research on GMM auto graders and language generation tasks such as LSTM, LSTM-CRF, LSTM-seq2seq, and GLU. The aim of this project is to bring an open framework for evaluating the performance of large training sets, making it easy and simple for researchers to evaluate existing and new deep-learning models.

The direct use of the model involves text generation, exploring characteristics of language generated by a language model such as Cloze tests, counterfactuals, and generations with reframings. The downstream usage covers tasks that leverage language models including Information Extraction, Question Answering, and Summarization. However, one should also check out the BLOOM License to be wary of the restrained use so that such frivolous usage of the software is avoided at all costs and any violations or deviations from the allowed usage can be monitored and reported. These guidelines are vital to prevent the abuse of the platform by any spammy/harassing/deceptive content creation.

Final Words: The Road Ahead for BLOOM

Now that the Hugging Face team has created a powerful model with T0++, it’s time to unlock its potential. We’re starting by exploring ways to make the model easier to use and experiment with, then adding new languages and models along with more capabilities and control.

Over time, the team plans to BLOOM project further to increase the number of models and languages it supports, support for additional architecture types, better performance metrics for tuning models and languages, and more sophisticated methods for writing computer programs in specialized languages. In the meantime, there’s a lot to discover with what has been written so far, so the Hugging Face teams want to help our community continue to grow its use of BLOOM by providing some key educational resources.

This has been a wonderful and inspiring start to the BLOOM movement, and the Hugging Face team is eager to see what the future holds. They encourage all the open source community members to continue practicing their skills on their own, as well as participating in grassroots efforts to keep the BLOOM family strong.

This is only the beginning for BLOOM and the Hugging Face team looks forward to continuing to work together with other communities and organizations interested in exploring how this model might impact their community on a local level, besides having a substantial and evident global impact. Now that our model has been validated, the Hugging Face team is ready to support community efforts to expand it and is optimistic about the open-source community participation in fueling grassroots efforts to keep the BLOOM family strong and to blossom the evolution of the BLOOM model.

Priya has about 7 years of experience in Market Research. Currently, she is working for Valasys Media, as a content writer, which is amongst the top B2B Media Publishers across the globe. She has been preparing several personalized reports for our clients & has done a lot of research on market segmentation, cluster analysis of audiences & inbound methodologies. She has worked with government institutes as well as corporate houses in several projects. She possesses various interests and believes in a data-driven approach to problem solving. She holds a post-graduation in science also writes extensively on all things about life besides marketing, science, data science and statistics. She is a firm believer in higher realities and that there’s always more to life than we understand. She is a psychic healer and a tarot practitioner, who believes in a spiritual way of living and practices Yoga and meditation. When not writing you can find her enjoying music or cooking. 

Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.