The Sequence AI of the Week #903: Laguna, the 118 Billion Parameters that Walks Into a Trillion-Parameter Bar
Substack Jesus Rodriguez ● Covered by 3 sources
Poolside's new coding AI, Laguna S 2.1, has just 118B parameters but beats models 13x its size on tough benchmarks. It's proof that scaling up isn't the only path to smarter AI—training approach matters more than raw size.
Based on reporting by Substack, Jesus Rodriguez — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Plot every open-weight model's parameter count against its Terminal-Bench 2.1 score and you get the graph everyone expects: a smooth upward slope, bigger models doing better, no surprises. Then Laguna S 2.1 shows up at 118 billion parameters with a score of 70.2%, sitting well above where that trend line says it should be.
For context, DeepSeek-V4-Pro-Max weighs in at 1.6 trillion parameters — more than 13 times the size — and only manages 64.0. Inkling, at 975 billion parameters, scores 63.8. Nemotron 3 Ultra, at 550 billion, lands at 56.4. So the smallest model in this comparison is beating models that are 5x, 8x, and 13x its size. On DeepSWE, a benchmark that's tougher and less picked-over, the gap turns into a chasm: Laguna hits 40.4 while DeepSeek-V4-Pro-Max manages just 9.0.
A result like this — huge score advantage, tiny fraction of the parameters — is usually the point where you start hunting for a broken benchmark or a contaminated training set. Poolside, the company behind Laguna, seems to have expected that skepticism. They published every trajectory from every trial in the final evaluation run. Anyone can go look at exactly what the model did on each task, not just the aggregate number.
That transparency choice says a lot about how this release was built. It's not a company trying to sneak past scrutiny with a flashy chart. It's a company betting that if you show your work, the numbers will hold up on their own. Poolside is not a household name yet, but results like these — especially paired with open trajectories instead of just marketing slides — tend to change that fast.
My take — AI-written commentary, not fact-checked reporting
I'll take a transparent 118B model that shows its receipts over a trillion-parameter black box any day. The industry's obsession with parameter count as a proxy for capability has always been lazy, and Laguna is a pretty loud reminder that training quality and evaluation design can matter more than brute-force scale. If more labs published full trajectories instead of cherry-picked benchmark bars, we'd have a lot fewer inflated claims to sort through.
Read more about this at: Substack
Related stories
[AINews] "Laguna S 2.1 Released: Cheaper than Deepseek v4 Flash, Better than V4 Pro"
Latent Space · 1 month ago ·
28
Poolside Releases Laguna S 2.1, an Open-Weight Agentic Coding Model Punching Above Its Weight Class on SWE-Bench Multilingual
MarkTechPost · 1 month ago ·
6