TLDRocket
Sign in

Cosmopedia: how to create large-scale synthetic data for pre-training Large Language Models

Hugging Face Blog

Researchers at Hugging Face created Cosmopedia, an open synthetic dataset containing 25 billion tokens generated using Mixtral-8x7B to replicate the training data behind Microsoft's Phi-1.5 language model. The dataset comprises over 30 million files across textbooks, blog posts, stories, and WikiHow articles, with less than 1% duplicate content achieved through extensive prompt engineering across 145 web-clustered topics and curated educational sources. The release includes the generation code, the full dataset, and a 1-billion-parameter model trained on it, enabling the community to reproduce high-performance language model training without proprietary data or models.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.