Will It Mythos?
TLDR Dev
A researcher built a benchmark to test whether Anthropic's Mythos security vulnerability detector is uniquely capable compared to other AI models, using nine confirmed bugs that Mythos had previously found. The benchmark tested 40+ models by asking them to identify and describe bugs in real code repositories without hints, with results showing no model performed perfectly and costs ranging from under $1 to over $100 per test run. The findings indicate that while Mythos appears strong, several cheaper models like Qwen 3.6 and DeepSeek are competitively capable, suggesting Mythos's advantage may not be as unique as Anthropic's security restrictions imply.
Why it matters
Mythos is a security tool believed to be amazing at finding vulnerabilities, but there's skepticism about the validity of its claims. A benchmarking project was launched to test whether other AI models could match Mythos's performance in identifying security bugs, with preliminary results showing that while some models performed surprisingly well, none consistently outperformed Mythos.