Thinking Machines Lab Releases Inkling-Small: A 276B Total, 12B Active Open Weights Multimodal MoE Model
MarkTechPost Asif Razzaq ● Covered by 33 sources
Thinking Machines Lab released Inkling-Small, an open weights multimodal Mixture-of-Experts model with 276B total parameters and 12B active under Apache 2.0 license. The NVFP4 quantized checkpoint requires 180GB of aggregated VRAM, deployable on a single NVIDIA B300 GPU or two H200s, making it accessible to startups and mid-size enterprises. The smaller model surpasses its 975B-parameter teacher Inkling on reasoning and coding benchmarks including SWE-bench Verified (80.2% vs 77.6%) and ARC-AGI-2 (40.1% vs 36.5%), while regressing on factual recall tasks.
Why it matters
Inkling-Small matches Inkling at a quarter the size, and its NVFP4 checkpoint runs on one NVIDIA B300 GPU The post Thinking Machines Lab Releases Inkling-Small: A 276B Total, 12B Active Open Weights Multimodal MoE Model appeared first on MarkTechPost.
Related stories
Thinking Machines released Inkling, an open-weight customizable multimodal model
The Neuron · 2 weeks ago ·
50
Inkling: Our open-weights model
Simon Willison · 2 weeks ago ·
39