Terminal-Bench 2.1
Benchmark ● Covered in 7 stories + Follow
Terminal-Bench 2.1 is a coding and agent benchmark used to evaluate long-horizon, multi-step software engineering performance. In recent coverage, multiple model releases reported Terminal-Bench 2.1 results, including Google’s Gemini 3.8 Flash (89.4%), DeepSeek’s V4-Flash-0731 (82.7, with independent testing noting 79%), and Meta’s Muse Code beta (evaluated on 89 tasks). The benchmark is also referenced as a metric for comparing agentic coding systems and guiding availability decisions across model providers.
Updated 12 September 2026
Latest developments
Z.ai Releases GLM-5.3-Flash: A 320B-A18B Natively Multimodal MoE With a 1M-Token Context
MarkTechPost · 3 weeks ago ·
4
IBM Releases Granite 4.2: Bringing Native Reasoning and Agentic RL to Open Enterprise Models
MarkTechPost · 3 weeks ago ·
12
Pokee AI Releases Pokee-Isaac 28B: A 10M-Token Context Agentic Model Built to Run Inside the Customer Boundary
MarkTechPost · 1 month ago ·
9
Meta AI Releases Muse Code (Beta): A Terminal Coding Agent Powered by the New Muse Spark 1.2 Model
MarkTechPost · 1 month ago ·
7
DeepSeek’s smaller model just outperformed its own flagship
The New Stack · 1 month ago ·
29
Gemini 3.5: frontier intelligence with action
Google DeepMind · 4 months ago ·
18
September 2026
Google launches Gemini 3.8 Flash and related agentic video understanding capabilities for its Gemini Flash models Model release
August 2026
Z.ai released and open-sourced GLM-5.3-Flash, a natively multimodal mixture-of-experts GLM model with a 1M-token context Open source release
Meta releases Muse Code terminal agent and Muse Spark 1.2 coding model Model release
- Z.ai Releases GLM-5.3-Flash: A 320B-A18B Natively Multimodal MoE With a 1M-Token Context
- IBM Releases Granite 4.2: Bringing Native Reasoning and Agentic RL to Open Enterprise Models
- Pokee AI Releases Pokee-Isaac 28B: A 10M-Token Context Agentic Model Built to Run Inside the Customer Boundary
- Meta AI Releases Muse Code (Beta): A Terminal Coding Agent Powered by the New Muse Spark 1.2 Model
- DeepSeek’s smaller model just outperformed its own flagship
May 2026
Google releases Gemini 3.5 Flash and Gemini Omni models with enhanced agentic capabilities and video generation Model release
Relationships
Products & technology
- Gemini 3.8 Flash integrated with this benchmark · 1 source
- Granite 4.2 derived from this benchmark · 1 source