Poolside AI releases Laguna 118B, an open-weight model demonstrating exceptional efficiency with significant performance gains over much larger competitors
Open source release ● Confirmed 92% confidence first seen
Poolside AI released Laguna S2.1, an open-weight model with 118 billion parameters that activates 8 billion per token and supports one-million-token context windows. The model substantially outperforms much larger models like DeepSeek-V4-Pro-Max on multiple benchmarks, achieving 70.2% on Terminal-Bench 2.1 and 40.4 on DeepSWE despite using 13 times fewer parameters, challenging conventional assumptions about the relationship between model scale and performance.
Decision brief
- What changed
- Poolside AI released Laguna S2.1, a 118-billion-parameter open-weight model (8B active per token, 1M-token context) that outperforms the much larger DeepSeek-V4-Pro-Max (1.6T parameters) on coding-related benchmarks, scoring 70.2% vs 64.0% on Terminal-Bench 2.1 and 40.4 vs 9.0 on DeepSWE.
- Why it matters
- This suggests architecture and training efficiency, not just raw parameter count, can drive frontier-level coding performance, potentially lowering the compute and infrastructure cost required to field competitive open-weight models. For enterprises evaluating AI vendors or building internal capabilities, this changes the calculus around whether to bet on massive proprietary models versus smaller, cheaper, open-weight alternatives for coding and agentic tasks.
- Evidence
- Two independent TheSequence newsletter issues (#901 and #903) report the release and benchmark figures consistently, with #903 providing detailed head-to-head comparisons against DeepSeek-V4-Pro-Max; a third source (Exponential View) covers unrelated market data and does not corroborate or contradict the Laguna claims.
- What remains uncertain
- The benchmarks (Terminal-Bench 2.1, DeepSWE) are narrow and coding/agentic-task specific, so it's unclear how Laguna performs on broader reasoning, safety, or production reliability metrics; independent third-party verification beyond the TheSequence reporting is not present in this coverage, and DeepSeek-V4-Pro-Max's benchmark scores are cited only via the same source.
- Monitor next
- Watch for independent third-party benchmark reproductions or enterprise adoption reports of Laguna S2.1 in production coding/agentic workflows over the coming weeks.
Analytical support, not advice — assumptions and open questions stated above.