TLDRocket
Sign in

Reflection AI Introduces Beam: A 501B Open-Weight MoE Model With 23B Active Parameters for Coding and Agentic Workloads

MarkTechPost Asif Razzaq ● Covered by 3 sources

Reflection AI unveiled Beam, a 501B open-weight model built for coding and agent work. It’s still not self-hostable, but it claims less compute than bigger rivals.

Based on reporting by MarkTechPost, Asif Razzaq — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Reflection AI has put Beam forward as its first open-weight model, and it’s not trying to win by brute force. Beam is a sparse mixture-of-experts system with 501B total parameters, but only 23B are active per token. The target is clear: coding, reasoning, and agentic work, with Reflection saying it can compete with larger open models while using 3 to 4 times less inference compute on reasoning benchmarks.

There’s a catch, and it matters. Beam is not ready for self-hosting yet; it’s in final red-teaming, and early access is being handled through a waitlist on Reflection’s own platform. The company also says Kimi K3 still leads on raw capability, so Beam’s main pitch is efficiency at inference time rather than a clean win on every task. Users can tune a reasoning-effort parameter, dialing it down for shorter answers or up for harder problems that need more thought.

The pretraining recipe is unusually aggressive. Beam was trained from scratch on 23.8 trillion tokens drawn from the web, public sources, and proprietary licensed data, with Reflection saying about 95% of raw internet tokens were removed during curation. The team says it preserved roughly 1.8 trillion high-quality tokens that normal filters would have thrown away. It also used an architecture that mixes local and global attention with routed experts, and claims the busiest expert ended pretraining at just 1.04x average load.

The compute story is just as heavy. Pretraining finished in under 4 weeks on 6,144 NVIDIA GB300 NVL72 GPUs, with goodput reaching 92.3% near the end after 9 semi-automatic rewinds. Midtraining pushed the effective context to 1M tokens. Then came the reinforcement-learning phase: 10.5K NVIDIA GB300 GPUs, 4 weeks, more than 100 million rollouts, and about 1.3 billion sandboxes across nearly 1 million coding, agentic, and STEM environments.

Reflection says the RL system stayed stable even with one-day staleness and 107 weight versions of lag. It also reports 110K concurrent rollouts on average, a median of about 12 seconds for new weights to reach inference, and 71 inference incidents handled without stopping training. On its own numbers, Beam scores 80.9 on SWE-bench Verified and 80.1 on Terminal Bench v2.1, which puts it close to the front of the open-weight pack but not at the top everywhere. Apache 2.0 weights are planned for later in October 2026.

My take — AI-written commentary, not fact-checked reporting

Open-weight AI keeps drifting toward a simple split: some teams chase raw size, others chase efficiency and shipping. Reflection is clearly betting that fewer active parameters and better training plumbing can beat a bigger pile of silicon, which is a much healthier obsession than chasing leaderboard confetti. The annoying part is that the model is still behind final red-teaming, so the industry gets another glossy preview before the thing is actually in the wild.

Read more about this at: MarkTechPost

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.