Poolside AI releases Laguna S 2.1, an open-weight 118B-parameter coding model with sparse mixture-of-experts architecture
Open source release ● Confirmed 95% confidence first seen
Poolside AI released Laguna S 2.1, a 118-billion-parameter open-weight mixture-of-experts model designed for coding tasks and agent work, achieving 78.5% on SWE-Bench Multilingual and 70.2% on Terminal-Bench 2.1 while activating only 8 billion parameters per token. The model reportedly outperforms much larger competitors like DeepSeek V4 Pro-Max while remaining efficient enough to run on local hardware with 4-bit quantization. The release demonstrates Poolside's systematic engineering approach that enables rapid model development cycles completing in under nine weeks.
Decision brief
- What changed
- Poolside AI released Laguna S 2.1, a 118B-parameter open-weight sparse mixture-of-experts coding model that activates only 8B parameters per token, scoring 78.5% on SWE-Bench Multilingual and 70.2% on Terminal-Bench 2.1, and can run locally with 4-bit quantization on a single NVIDIA DGX Spark.
- Why it matters
- An open-weight model that matches or beats much larger closed and open competitors on coding benchmarks while running on local hardware lowers the cost and infrastructure barrier for deploying agentic coding tools in-house, reducing dependency on API-based providers. Poolside's claimed sub-nine-week model development cycle (via its 'Model Factory' approach) also signals a faster competitive cadence in the open-weight coding model space that could compress the shelf life of current vendor advantages.
- Evidence
- Benchmark scores and architecture details (118B total/8B active parameters, 78.5% SWE-Bench Multilingual, 70.2% Terminal-Bench 2.1) are consistently reported across all four sources, including MarkTechPost, TLDR Dev, and two Latent Space pieces; the Model Factory development claims come primarily from a single Latent Space interview with co-founder Eiso Kant.
- What remains uncertain
- Benchmark comparisons against DeepSeek V4 Pro-Max and Nemotron 3 Ultra rely on Poolside's own reported evaluations rather than independent third-party verification, and DeepSWE v1.1 scores (40.4%) suggest performance may vary by benchmark. The sustainability and reproducibility of the '8-week model cycle' and '10,000-20,000 experiments monthly' claims are not independently confirmed outside the company's own account.
- Monitor next
- Watch for independent third-party benchmark reproductions or enterprise adoption reports confirming whether Laguna S 2.1's performance and cost advantages hold up outside Poolside's own published evaluations.
Analytical support, not advice — assumptions and open questions stated above.