TLDRocket
Sign in

Can an Open Model Do Security Research? Cantina’s apex-flash-1 Solves 40 of 60 Held-Out Bug Tasks

MarkTechPost Michal Sutter

Cantina’s new open security model solved 40 of 60 bug tasks. It’s cheaper than Claude Opus 5 High and can run locally, if you’ve got the GPU muscle.

Based on reporting by MarkTechPost, Michal Sutter — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Cantina Security and Yeta Labs have put out apex-flash-1, an open-weights model tuned for vulnerability research rather than generic chat. It’s built by fine-tuning Z.ai’s GLM-5.3-Flash with reinforcement learning, and the weights are on Hugging Face under the MIT license. That makes it unusually practical for a security model: it’s meant to be run, not just admired from a slide deck.

The model is huge. Cantina says the safetensors metadata puts it at 321.3 billion total parameters, while the GLM-5.3-Flash base is a mixture-of-experts model with 18 billion active parameters. Training used GRPO, a rank-256 LoRA, and selective full-parameter training. The data itself came from 50 real vulnerability cases turned into 150 tasks, with each case split into guided whitebox, focused whitebox, and focused blackbox variants.

The mix of bugs says a lot about what the model was taught to care about. Authorization, identity and scope issues account for 72% of the cases. Accounting and numerical precision make up another 18%. The rest cover time validation, business rules and SSRF. Cantina says the reinforcement learning rollouts ran inside the Codex agent harness on production-like software and protocol environments.

On its internal benchmark, apex-flash-1 solved 40 of 60 held-out tasks, which is 66.7% pass@1. Cantina says that run cost about $2.38. The base GLM-5.3-Flash model solved 36 of 60 for about $4.56, while Claude Opus 5 High solved 43 of 60 for about $74.68. So Opus found three more tasks, but at roughly 31 times the cost per run. On Cantina’s math, that’s about $0.06 per solved task for apex-flash-1 versus $1.74 for Opus.

The catch is obvious. BF16 deployment needs about 640 GB of GPU memory, so this is not exactly a laptop hobbyist’s afternoon project. But Cantina’s pitch is aimed at defenders who want a capable worker model they can keep local and under control, rather than handing security work to a black-box service.

My take — AI-written commentary, not fact-checked reporting

This is the right kind of open-model story: not bigger chat, but a model aimed at a narrow job and judged on actual bug tasks. The closed-model crowd still gets to brag about a few more solves, but local control and a sane bill matter when the task is security research. Also, 640 GB of GPU memory is a lovely reminder that “open” is sometimes just a licensing word with a very expensive gym membership.

Read more about this at: MarkTechPost

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.