TLDRocket
Sign in

Thinking Machines Lab Releases Inkling: A 975B-Parameter Open-Weights Multimodal MoE With 41B Active Parameters And Controllable Thinking Effort

MarkTechPost Asif Razzaq Covered by 4 sources

Thinking Machines Lab released Inkling, a 975-billion-parameter open-weights multimodal mixture-of-experts model with 41 billion active parameters and a 1-million-token context window. The model was trained on 45 trillion tokens of text, images, audio, and video, and features a controllable thinking-effort mechanism that allows users to trade inference cost for performance, achieving Terminal Bench 2.1 parity with Nemotron 3 Ultra using one-third the tokens. Inkling is available on Hugging Face and hosted platforms, enabling deployment of cost-tuned agentic systems and multimodal applications across voice, vision, and text inputs.

Why it matters

Thinking Machines Lab released Inkling on July 15, 2026, its first model trained from scratch. The full weights ship under Apache 2.0. It is a 975B-parameter Mixture-of-Experts transformer with 41B active parameters, a 1M-token context window, and native text, image, and audio input. The lab states plainly that Inkling is not the strongest model available, open or closed. It is positioned instead as a customization base, with controllable thinking effort as the practical differentiator. The post Thinking Machines Lab Releases Inkling: A 975B-Parameter Open-Weights Multimodal MoE With 41B Active Parameters And Controllable Thinking Effort appeared first on MarkTechPost.

Also covered by

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.