JetBrains Releases Mellum2.1: A 12B MoE Open Model for Coding Agents
MarkTechPost Asif Razzaq
JetBrains released Mellum2.1, an open 12B model for coding agents. It runs with only 2.5B active params and can self-host on your own GPUs.
Based on reporting by MarkTechPost, Asif Razzaq — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
JetBrains has put out Mellum2.1, the latest version of its Mellum line, and the pitch is refreshingly specific: an open model built for coding agents, fast sub-agents, and private deployment. It ships under Apache 2.0 on Hugging Face, keeps the same architecture as Mellum2, and gets its upgrade mostly from reinforcement learning in real software environments.
That matters because Mellum2.1 is not trying to win by brute force. It is a 12B mixture-of-experts model with 2.5B active parameters per token, 64 experts, and 8 active at a time. JetBrains says it has a 131,072-token context, 28 layers, grouped-query attention, and a vocabulary of 98,304 tokens. The weights are in bfloat16, and the model is meant to work as a reasoning system that emits its chain of thought before answering.
The training story is where JetBrains put most of its effort. RL moved from a short final stage to the main event, with tasks spanning math, competitive programming, science, tool use, and software engineering. For the coding part, the model works inside real repositories with shell and file-editing tools, gets rewarded when tests pass, and was trained across millions of sandboxes in thousands of environments. JetBrains also says it filtered open RL datasets to remove broken tests, unverifiable answers, and tasks that were too easy or impossible.
On JetBrains’ own benchmark pipeline, Mellum2.1 makes a sharp jump over Mellum2 on agentic coding. SWE-bench Verified rises from 2.0 to 47.0, SWE-bench Pro from 0.0 to 28.0, and Terminal-Bench 2.1 from 0.6 to 17.4. It also leads the group on LiveCodeBench v6 at 82.0, HumanEval+ at 91.5, MBPP+ at 79.4, and BFCL v4 at 62.3. Qwen3.5-9B still comes out ahead on SWE-bench Verified, SWE-bench Pro, AIME 25/26, and GPQA Diamond.
JetBrains says the architecture was left untouched, so speed stays close to Mellum2. On one NVIDIA H200 under heavy load, it serves almost twice as many tokens as Qwen3.5-9B, and for a single request the multi-token prediction setup makes it about 1.6x faster. Local users get a practical route too: the full model runs on vLLM or SGLang, and GGUF builds start at 7.0 GB for compact setups like llama.cpp, Ollama, and LM Studio.
My take — AI-written commentary, not fact-checked reporting
This is the kind of open-model release that actually matters: not a bigger number for a demo slide, but a model tuned for work people do. JetBrains is quietly making the case that coding agents should be local, inspectable, and not rented by the hour from someone else’s cloud. A rare sensible move in a field that loves to confuse scale with progress.
Read more about this at: MarkTechPost