TLDRocket
Sign in

Now in Nature: Retrofitting language models to operate over bytes

Allen Institute (AI2)

AI2 put its byte-level Bolmo research in Nature and opened new checkpoints. It’s a bet that models built from bytes can be both open and competitive, not just clever.

Based on reporting by Allen Institute (AI2) — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Allen Institute for AI says the research behind Bolmo has now been published in Nature, alongside new checkpoints on Hugging Face. The release pushes the same byte-level approach beyond the Olmo models that Bolmo started from, and it also includes Stage 1 checkpoints for researchers who want a quicker place to begin experimenting with the architecture.

The basic idea is simple enough. Most language models chop text into subwords, which usually works well but can stumble on spelling, rare words, unusual strings, or text that doesn’t fit neatly into a fixed vocabulary. Bolmo works lower down the stack, directly on the bytes that make up text — letters, punctuation, symbols, and everything else computers use to represent characters.

That shift buys flexibility. AI2 says byte-level models can better capture fine-grained structure such as whitespace and different writing systems without being tied to a fixed vocabulary. And because bytes are the foundation for digital data more broadly, the same approach could one day be useful beyond text, including for images and audio.

The hard part has always been performance. Training byte-level models from scratch has been expensive, which made them hard to keep up with as subword models kept improving. AI2’s answer is a process it calls byteifying: start with an already capable subword model, then convert it into a byte-level one with a relatively short extra training run.

That method produced Bolmo 1B and Bolmo 7B from AI2’s open Olmo models, and the lab says they were the first fully open byte-level language models competitive with strong subword models across a broad range of tasks. The Nature paper shows the trick works beyond Olmo too. AI2 applied it to Qwen 3 8B and Llama 3 8B, creating Bwen 8B and Blama 8B. Both come close to the models they were derived from, and Bwen 8B is described as the strongest byteified model yet, ahead of Bolmo 7B on AI2’s aggregate evaluation suite.

People are already poking at Bolmo for niche uses, including a Bolmo 7B fine-tune for Russian-to-English poetry and lyric translation. That fits AI2’s broader pitch: publish the model, publish the recipe, let other researchers test the assumptions, and maybe stop pretending one representation of text is the only one that matters.

My take — AI-written commentary, not fact-checked reporting

Open models only matter when they’re actually usable, and AI2 seems to get that. A shiny paper without checkpoints is just academic fan art; publishing the recipe, the weights, and the training path is the part that lets other people do real work. The byte-level bet is also refreshingly unglamorous, which is probably why it might age better than half the hype parade around AI this year.

Read more about this at: Allen Institute (AI2)

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.