TLDRocket
Sign in

The Technology Behind BLOOM Training

Hugging Face Blog Covered by 2 sources

Researchers trained BLOOM, a 176 billion parameter language model, using 384 NVIDIA A100 GPUs on France's Jean Zay supercomputer with a custom Megatron-DeepSpeed software framework. The training took 3.5 months to complete between March and July 2022, processing 350 billion tokens across 59 languages. The project demonstrated that large-scale multilingual model training could be accomplished by a distributed team using open-source software and published infrastructure, making the approach and final model publicly accessible.

Also covered by

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.