TLDRocket
Sign in

Training mRNA Language Models Across 25 Species for $165

Hugging Face

OpenMed built an end-to-end protein engineering pipeline combining structure prediction, sequence design, and codon optimization across 25 species for $165 in compute costs. CodonRoBERTa-large-v2 achieved a perplexity of 4.10 and CAI correlation of 0.40, outperforming ModernBERT by 6x, with four production models trained in 55 GPU-hours. The system enables researchers to move from a protein concept to synthesis-ready DNA sequences by applying transformer models trained specifically on biological codon statistics rather than natural language patterns.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.