Can an Open Model Do Security Research? Cantina’s apex-flash-1 Solves 40 of 60 Held-Out Bug Tasks
MarkTechPost Michal Sutter
Cantina’s new open security model solved 40 of 60 bug tasks. It’s cheaper than Claude Opus 5 High and can run locally, if you’ve got the GPU muscle.
Based on reporting by MarkTechPost, Michal Sutter — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Cantina Security and Yeta Labs have put out apex-flash-1, an open-weights model tuned for vulnerability research rather than generic chat. It’s built by fine-tuning Z.ai’s GLM-5.3-Flash with reinforcement learning, and the weights are on Hugging Face under the MIT license. That makes it unusually practical for a security model: it’s meant to be run, not just admired from a slide deck.
The model is huge. Cantina says the safetensors metadata puts it at 321.3 billion total parameters, while the GLM-5.3-Flash base is a mixture-of-experts model with 18 billion active parameters. Training used GRPO, a rank-256 LoRA, and selective full-parameter training. The data itself came from 50 real vulnerability cases turned into 150 tasks, with each case split into guided whitebox, focused whitebox, and focused blackbox variants.
The mix of bugs says a lot about what the model was taught to care about. Authorization, identity and scope issues account for 72% of the cases. Accounting and numerical precision make up another 18%. The rest cover time validation, business rules and SSRF. Cantina says the reinforcement learning rollouts ran inside the Codex agent harness on production-like software and protocol environments.
On its internal benchmark, apex-flash-1 solved 40 of 60 held-out tasks, which is 66.7% pass@1. Cantina says that run cost about $2.38. The base GLM-5.3-Flash model solved 36 of 60 for about $4.56, while Claude Opus 5 High solved 43 of 60 for about $74.68. So Opus found three more tasks, but at roughly 31 times the cost per run. On Cantina’s math, that’s about $0.06 per solved task for apex-flash-1 versus $1.74 for Opus.
The catch is obvious. BF16 deployment needs about 640 GB of GPU memory, so this is not exactly a laptop hobbyist’s afternoon project. But Cantina’s pitch is aimed at defenders who want a capable worker model they can keep local and under control, rather than handing security work to a black-box service.
My take — AI-written commentary, not fact-checked reporting
This is the right kind of open-model story: not bigger chat, but a model aimed at a narrow job and judged on actual bug tasks. The closed-model crowd still gets to brag about a few more solves, but local control and a sane bill matter when the task is security research. Also, 640 GB of GPU memory is a lovely reminder that “open” is sometimes just a licensing word with a very expensive gym membership.
Read more about this at: MarkTechPost
Related stories
Cisco Foundation AI Releases Antares: 350M and 1B Open-Weight Models That Localize Known Vulnerabilities Inside Real Codebases
MarkTechPost · 2 months ago ·
25
We have Mythos at Home: GLM 5.2 beats Claude in our Cyber Benchmarks
Semgrep · 3 months ago ·
11
The flaw-hunting machine that isn’t allowed to trust itself
Tech Funding News · 3 weeks ago ·
34