Thinking Machines Lab releases Inkling, a 975B-parameter open-weights multimodal mixture-of-experts model
Open source release ● Confirmed 95% confidence first seen
Thinking Machines Lab released Inkling, an open-weights multimodal mixture-of-experts model with 975 billion total parameters and 41 billion active parameters. The model features a 1-million-token context window, was trained on 45 trillion tokens of multimodal data, and includes a controllable thinking-effort mechanism for inference cost-performance tradeoffs. Inkling is positioned as a customizable alternative to proprietary models and is available on Hugging Face and the Tinker fine-tuning platform.
Decision brief
- What changed
- Thinking Machines Lab released Inkling, an open-weights multimodal mixture-of-experts model with 975B total parameters (41B active), a 1M-token context window, trained on 45T multimodal tokens, and a controllable thinking-effort mechanism; it's available on Hugging Face and the Tinker fine-tuning platform.
- Why it matters
- This gives enterprises a self-hostable, customizable alternative to proprietary APIs (OpenAI, Anthropic, etc.) with multimodal capability and tunable inference cost, reinforcing a broader shift of workloads toward open-weight models. Leaders evaluating build-vs-buy AI infrastructure decisions now have a credible large-scale open option that reportedly matches or approaches frontier proprietary performance at lower token cost, affecting vendor negotiation leverage and internal fine-tuning strategy.
- Evidence
- Four independent outlets (MarkTechPost, The Neuron, Ben's Bites, Deep Learning Weekly) consistently report the same core specs (975B/41B params, 1M context, 45T training tokens, controllable thinking effort) and availability on Hugging Face/Tinker, indicating reliable technical detail, though benchmark comparisons (e.g., Terminal Bench 2.1 parity with Nemotron 3 Ultra, positioning between Kimi 2.5 and 2.6) come from single-source claims not independently verified across all outlets.
- What remains uncertain
- Benchmark performance claims (parity with Nemotron 3 Ultra using one-third tokens, ranking between Kimi 2.5/2.6) rely on the releasing lab's or a single outlet's framing and lack independent third-party verification; actual enterprise deployment costs, safety/security posture, and real-world fine-tuning performance versus proprietary models remain unverified. It's also unclear how licensing terms or compute requirements affect practical accessibility for smaller organizations.
- Monitor next
- Watch for independent third-party benchmark evaluations and early enterprise deployment reports (e.g., fine-tuning results on Tinker) that test whether Inkling's claimed cost-performance tradeoffs hold up outside Thinking Machines Lab's own comparisons.
Analytical support, not advice — assumptions and open questions stated above.