Latest open artifacts (#21): Open model bonanza! Gemma 4, DeepSeek V4, Kimi K2.6, MiMo 2.5, GLM-5.1 & others. On CAISI's V4 assessment.
Interconnects Florian Brand
Multiple open-source AI labs released new frontier models this month, including DeepSeek V4, Gemma 4, Kimi K2.6, and others, prompting a fresh evaluation of how open models compare to closed alternatives. The Center for AI Standards and Innovation found open models lag by 3-7 months behind American frontier models, with the gap widening when measured against proprietary benchmarks like CTF-Archive-Diamond and PortBench. The benchmarking methodology itself remains contested, as standardized evaluation setups do not match how models are actually deployed—for example, using basic coding harnesses rather than the specialized environments models are trained for.
Why it matters
An eventful month with one flagship release after another
Related stories
Import AI 465: Open vs closed gaps; Kimi K3; Demis' big policy plan
Import AI · 1 month ago ·
20
Kimi K3 Redraws the Open Frontier, Muse Spark 1.1 Undercuts Competitors, Cloudflare Moves to Cut Off Crawlers
The Batch ·
32