TLDRocket
Sign in

T5Gemma: A new collection of encoder-decoder Gemma models

Google DeepMind Covered by 4 sources

Google introduced T5Gemma, a collection of encoder-decoder language models created by converting pretrained decoder-only Gemma 2 models into encoder-decoder architectures through a technique called adaptation. T5Gemma models range from 2B to XL sizes, with a 9B-2B variant that pairs a large encoder with a small decoder to optimize quality-efficiency trade-offs. The adapted models outperform their decoder-only equivalents on reasoning tasks—for example, T5Gemma 2B-2B instruction-tuned achieves a 12-point improvement on MMLU over Gemma 2 2B—and the pretrained and fine-tuned checkpoints are now available on Hugging Face and Kaggle.

Why it matters

Introducing T5Gemma, a new collection of encoder-decoder LLMs.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.