TLDRocket
Sign in

Reflection’s Beam: an open model you can run on your own servers with lower compute needs

Reflection ● Covered by 6 sources

Reflection launched Beam, a 501B open-weight model you can run on your own servers. It uses 23B active params, so it’s built to be cheaper at inference than its size sounds.

Based on reporting by Reflection — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Reflection has introduced Beam, its first open-weight model, and the pitch is straightforward: big model, smaller appetite. Beam is a sparse mixture-of-experts system with 501 billion total parameters but only 23 billion active at a time, aimed at coding, reasoning, and agentic work.

The company says that tradeoff came from brute-force pretraining and an unusually heavy reinforcement-learning push. Beam was pretrained on 23.8 trillion curated tokens from the web and licensed proprietary data, then pushed through more than 100 million rollouts across 10.5K NVIDIA GB300 GPUs over four weeks. Reflection says that training run used roughly 1.3 billion sandboxes and that capabilities kept improving as RL compute scaled up, with no sign of a plateau.

The interesting part is not just that Beam can do the work. It’s that Reflection is arguing it can do it with less inference compute than comparable open models. The company says Beam is competitive with models such as GLM 5.2 on coding and agentic tasks, approaches Qwen 3.8-Max in those areas, and matches GLM-5.2 on advanced reasoning while using 3–4× less inference compute. It also says Beam is less capable than Kimi K3 on raw capability, but more efficient at serving time.

That efficiency story extends into the training stack. Reflection built an asynchronous RL system with stable learning under policy staleness, fast weight updates, and enough infrastructure to keep an average of 110K rollouts running at once. It says new weights reached the inference fleet in about 12 seconds median, and that 71 inference incidents were handled without stopping training.

Beam is still in final red-teaming and evaluation, with release of the weights, technical report, model card, and developer artifacts promised later this month. Early access is open now. For teams that care less about bragging rights and more about cost per useful token, that’s the real hook.

My take — AI-written commentary, not fact-checked reporting

Open-weight models keep getting sold as freedom, but the practical win is simpler: lower serving costs and fewer cloud tears. Reflection is leaning into the part that matters most to companies, not demos. That’s the right pitch, and it’s also a reminder that the open race is increasingly about efficiency, not just raw parameter theater.

Read more about this at: Reflection

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.