TLDRocket
Sign in

Thinking Machines Lab releases Inkling, a 975B-parameter open-weights multimodal mixture-of-experts model

Open source release Confirmed 95% confidence first seen

Thinking Machines Lab released Inkling, an open-weights multimodal mixture-of-experts model with 975 billion total parameters and 41 billion active parameters. The model features a 1-million-token context window, was trained on 45 trillion tokens of multimodal data, and includes a controllable thinking-effort mechanism for inference cost-performance tradeoffs. Inkling is positioned as a customizable alternative to proprietary models and is available on Hugging Face and the Tinker fine-tuning platform.

Decision brief

What changed
Thinking Machines Lab released Inkling, an open-weights multimodal mixture-of-experts model with 975B total parameters (41B active), a 1M-token context window, trained on 45T multimodal tokens, and a controllable thinking-effort mechanism; it's available on Hugging Face and the Tinker fine-tuning platform.
Why it matters
This gives enterprises a self-hostable, customizable alternative to proprietary APIs (OpenAI, Anthropic, etc.) with multimodal capability and tunable inference cost, reinforcing a broader shift of workloads toward open-weight models. Leaders evaluating build-vs-buy AI infrastructure decisions now have a credible large-scale open option that reportedly matches or approaches frontier proprietary performance at lower token cost, affecting vendor negotiation leverage and internal fine-tuning strategy.
Affected roles
CEO CTO CISO CFO
Evidence
Four independent outlets (MarkTechPost, The Neuron, Ben's Bites, Deep Learning Weekly) consistently report the same core specs (975B/41B params, 1M context, 45T training tokens, controllable thinking effort) and availability on Hugging Face/Tinker, indicating reliable technical detail, though benchmark comparisons (e.g., Terminal Bench 2.1 parity with Nemotron 3 Ultra, positioning between Kimi 2.5 and 2.6) come from single-source claims not independently verified across all outlets.
What remains uncertain
Benchmark performance claims (parity with Nemotron 3 Ultra using one-third tokens, ranking between Kimi 2.5/2.6) rely on the releasing lab's or a single outlet's framing and lack independent third-party verification; actual enterprise deployment costs, safety/security posture, and real-world fine-tuning performance versus proprietary models remain unverified. It's also unclear how licensing terms or compute requirements affect practical accessibility for smaller organizations.
Monitor next
Watch for independent third-party benchmark evaluations and early enterprise deployment reports (e.g., fine-tuning results on Tinker) that test whether Inkling's claimed cost-performance tradeoffs hold up outside Thinking Machines Lab's own comparisons.

Analytical support, not advice — assumptions and open questions stated above.

Source coverage

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.