The Next Top Open Model, Google Voice Agents, DeepSeek Shrinks Caches: Plus a letter on open weight cybersecurity capabilities
The Batch Analytics DeepLearning.AI
Open-weight model GLM-5.3 was analyzed as approaching Claude Mythos’ cyber capability, including a 12% versus 14% success rate on ExploitBench tasks at comparable token counts and evidence that $20.40 of GLM-5.3-Flash tokens could find a recently disclosed Chrome flaw. Xiaomi then released MiMo-V2.6-Pro-RL, MiMo-V2.6-Flash, and a distilled variant, with MiMo-V2.6-Pro-RL costing $0.435 per million input tokens (and free downloads under MIT) while reporting RL training details and CyberBench v1.1 results like 75.36% accuracy for MiMo-V2.6-Flash. Defenders are pushed to accelerate sandboxing, monitoring, and vulnerability mitigation as attackers can access strong open-weight cyber-capable models and the cost to generate attacks is dropping.
Why it matters
The Batch News & Insights: Anthropic recently released an encouraging analysis of the cyber capabilities of open weight model GLM-5.3.
Related stories
Can an Open Model Do Security Research? Cantina’s apex-flash-1 Solves 40 of 60 Held-Out Bug Tasks
MarkTechPost · 3 days ago ·
8