Thinking Machines Lab Releases Inkling-Small: A 276B Total, 12B Active Open Weights Multimodal MoE Model
MarkTechPost 3 weeks ago 15 ● 33 sources
Thinking Machines Lab released Inkling-Small, an open weights multimodal Mixture-of-Experts model with 276B total parameters and 12B active under Apache 2.0 license. The NVFP4 quantized checkpoint requires 180GB of aggregated VRAM, deployable on a single NVIDIA B300 GPU or two H200s, making it accessible to startups and mid-size enterprises. The smaller model surpasses its 975B-parameter teacher Inkling on reasoning and coding benchmarks including SWE-bench Verified (80.2% vs 77.6%) and ARC-AGI-2 (40.1% vs 36.5%), while regressing on factual recall tasks.