TLDRocket
Sign in

OlmPool: How small architectural choices compound to undermine long context extension

Allen Institute (AI2)

Researchers found that four architectural choices in language models—QK normalization, grouped-query attention, sliding window attention, and shorter pretraining context length—individually have modest negative effects on long context performance but combine to reduce benchmark scores by up to 47%. The study used OlmPool, a suite of 26 7-billion-parameter models trained on 140 billion tokens with identical data but different architectures, to isolate these effects. The findings show that Llama 3's long context capabilities come primarily from architecture rather than training data, meaning context extension recipes validated on Llama may not transfer to other model families without adjustment.

Why it matters

OlmPool is a controlled suite of 26 models showing how small architecture choices can compound to make long-context extension much harder, even when training data and extension recipes are held constant.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.