Terminal-Bench 2.1
Benchmark ● Covered in 7 stories + Follow
Terminal-Bench 2.1 is a coding and agent benchmark used to evaluate long-horizon, multi-step software engineering performance. In recent coverage, multiple model releases reported Terminal-Bench 2.1 results, including Google’s Gemini 3.8 Flash (89.4%), DeepSeek’s V4-Flash-0731 (82.7, with independent testing noting 79%), and Meta’s Muse Code beta (evaluated on 89 tasks). The benchmark is also referenced as a metric for comparing agentic coding systems and guiding availability decisions across model providers.
Updated 12 September 2026
Latest developments
Z.ai Releases GLM-5.3-Flash: A 320B-A18B Natively Multimodal MoE With a 1M-Token Context
MarkTechPost · 3 weeks ago ·
4
IBM Releases Granite 4.2: Bringing Native Reasoning and Agentic RL to Open Enterprise Models
MarkTechPost · 3 weeks ago ·
12
Pokee AI Releases Pokee-Isaac 28B: A 10M-Token Context Agentic Model Built to Run Inside the Customer Boundary
MarkTechPost · 1 month ago ·
9
Meta AI Releases Muse Code (Beta): A Terminal Coding Agent Powered by the New Muse Spark 1.2 Model
MarkTechPost · 1 month ago ·
7
DeepSeek’s smaller model just outperformed its own flagship
The New Stack · 1 month ago ·
29
Gemini 3.5: frontier intelligence with action
Google DeepMind · 4 months ago ·
18
2026
Google launches Gemini 3.8 Flash and related agentic video understanding capabilities for its Gemini Flash models Model release
Z.ai released and open-sourced GLM-5.3-Flash, a natively multimodal mixture-of-experts GLM model with a 1M-token context Open source release
Meta releases Muse Code terminal agent and Muse Spark 1.2 coding model Model release
DeepSeek releases V4-Flash-0731 model with improved agentic and coding performance at reduced pricing Model release
Google releases Gemini 3.5 Flash and Gemini Omni models with enhanced agentic capabilities and video generation Model release
- Gemini 3.8 Flash
- Z.ai Releases GLM-5.3-Flash: A 320B-A18B Natively Multimodal MoE With a 1M-Token Context
- IBM Releases Granite 4.2: Bringing Native Reasoning and Agentic RL to Open Enterprise Models
- Pokee AI Releases Pokee-Isaac 28B: A 10M-Token Context Agentic Model Built to Run Inside the Customer Boundary
- Meta AI Releases Muse Code (Beta): A Terminal Coding Agent Powered by the New Muse Spark 1.2 Model
- DeepSeek’s smaller model just outperformed its own flagship
- Gemini 3.5: frontier intelligence with action
Relationships
Products & technology
- Gemini 3.8 Flash integrated with this benchmark · 1 source
- Granite 4.2 derived from this benchmark · 1 source