TLDRocket
Sign in

Thinking Machines Lab releases Inkling, a 975B-parameter open-weights multimodal mixture-of-experts model

Open source release Confirmed 95% confidence first seen

Thinking Machines Lab released Inkling, an open-weights multimodal mixture-of-experts model with 975 billion total parameters and 41 billion active parameters. The model features a 1-million-token context window, was trained on 45 trillion tokens of multimodal data, and includes a controllable thinking-effort mechanism for inference cost-performance tradeoffs. Inkling is positioned as a customizable alternative to proprietary models and is available on Hugging Face and the Tinker fine-tuning platform.

Decision brief

What changed
Thinking Machines Lab released Inkling, an open-weights multimodal mixture-of-experts model with 975B total parameters (41B active), a 1M-token context window, trained on 45T multimodal tokens, and a controllable thinking-effort mechanism; it's available on Hugging Face and the Tinker fine-tuning platform.
Why it matters
This gives enterprises a self-hostable, customizable alternative to proprietary APIs (OpenAI, Anthropic, etc.) with multimodal capability and tunable inference cost, reinforcing a broader shift of workloads toward open-weight models. Leaders evaluating build-vs-buy AI infrastructure decisions now have a credible large-scale open option that reportedly matches or approaches frontier proprietary performance at lower token cost, affecting vendor negotiation leverage and internal fine-tuning strategy.
Affected roles
CEO CTO CISO CFO
Evidence
Four independent outlets (MarkTechPost, The Neuron, Ben's Bites, Deep Learning Weekly) consistently report the same core specs (975B/41B params, 1M context, 45T training tokens, controllable thinking effort) and availability on Hugging Face/Tinker, indicating reliable technical detail, though benchmark comparisons (e.g., Terminal Bench 2.1 parity with Nemotron 3 Ultra, positioning between Kimi 2.5 and 2.6) come from single-source claims not independently verified across all outlets.
What remains uncertain
Benchmark performance claims (parity with Nemotron 3 Ultra using one-third tokens, ranking between Kimi 2.5/2.6) rely on the releasing lab's or a single outlet's framing and lack independent third-party verification; actual enterprise deployment costs, safety/security posture, and real-world fine-tuning performance versus proprietary models remain unverified. It's also unclear how licensing terms or compute requirements affect practical accessibility for smaller organizations.
Monitor next
Watch for independent third-party benchmark evaluations and early enterprise deployment reports (e.g., fine-tuning results on Tinker) that test whether Inkling's claimed cost-performance tradeoffs hold up outside Thinking Machines Lab's own comparisons.

Analytical support, not advice — assumptions and open questions stated above.

Source coverage

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.