TLDRocket
Sign in

How researchers adapted Dolma for better Thai language models

Allen Institute (AI2)

AI2’s open data tool Dolma helped Thai researchers build Mangosteen, a 47-billion-token training corpus. It had to be rewired for Thai, and that made the model better on local knowledge.

Based on reporting by Allen Institute (AI2) — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

AI2’s Dolma was built to be changed. Now a Thai research team has done exactly that, using the open data-curation toolkit as the base for Mangosteen, a 47-billion-token pretraining corpus for Thai language models.

The point wasn’t just speed. It was survival. Building a full data pipeline from scratch would have been too much for a small team, according to Wannaphong Phatthiyaphaibun, a project lead and PhD student at VISTEC. Dolma gave them a working starting point, so they could spend their effort on Thai-specific problems instead of rebuilding every piece of the workflow.

And Thai-specific problems turned out to be real. Some of the standard pipeline steps, especially sentence- and paragraph-level deduplication, wiped out nearly all the data because Thai doesn’t mark sentence boundaries the way English does. The team kept what still made sense, like document- and URL-level deduplication, then adjusted quality filters, swapped in different language tools, and added rules for patterns common in Thai web text.

They also found gaps in the existing Thai data ecosystem. Publicly available pretraining datasets had not been heavily audited by Thai speakers and leaned too much on web crawls. The researchers said some of that data was unsuitable, while important sources such as books, research papers, official websites, and YouTube subtitles were missing.

The result was a corpus that did more with less. In tests with Thai LLMs, the curation pipeline removed more than 80% of the Common Crawl data it started with and nearly half of FineWeb2, while keeping performance equal or better than models trained on the larger datasets. Those gains also showed up in bigger models, which performed better on Thai cultural knowledge evaluations. For the team, that was the proof point: local expertise plus open tooling can make models fit the communities they’re meant to serve.

My take — AI-written commentary, not fact-checked reporting

This is the part of open AI that matters: not free downloads, but editable machinery. If a pipeline falls apart on Thai sentence boundaries, the fix should come from Thai researchers, not from a committee with a gloss on multilinguality. Closed systems love to call that “handled”; open tools actually make it possible.

Read more about this at: Allen Institute (AI2)

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.