TLDRocket
Sign in

Microsoft unveils AI security tools it says outperform competing platforms

Ars Technica Dan Goodin Covered by 4 sources

Microsoft launched two AI security tools it says beat Anthropic, Google and OpenAI. They're cheaper too, but still in preview, untested at scale.

Based on reporting by Ars Technica, Dan Goodin — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Microsoft is making a big claim this week: its new MDASH tool, running on something called MAI-Cyber-1-Flash, scored 96 percent on CyberGYM, a benchmark used to test cybersecurity AI. That's 12 points ahead of Anthropic's Mythos, and Microsoft says it edges out Google Gemini and OpenAI GPT too. And the kicker is price — this version of MDASH costs half as much to run as the last one.

The second announcement, Project Perception, is a different kind of tool entirely. Rather than one model doing everything, it's a set of specialized AI agents split into red, blue and green teams — hunting for vulnerabilities, figuring out how dangerous they actually are, and then fixing them. The platform picks which model handles which job based on effectiveness and what it'll cost the customer, decisions Microsoft says come from ongoing research and benchmarking across both frontier and specialized models. The pitch is that Project Perception can handle 90 percent of security tasks more cheaply than rival platforms, meaning customers would only need to pay for something pricier on the remaining 10 percent.

Microsoft frames all of this as a response to a fundamental shift in the threat landscape — its words, not a paraphrase of some old playbook. AI, the company argues, is letting attackers move faster and at greater scale, while security teams are stuck stitching together fragments of data across systems that were never designed for this pace of assault.

But there's a glaring gap in Microsoft's pitch. Both tools are still in preview, and the company said nothing about the risks of turning security decisions over to AI agents — an omission that lands awkwardly just a week after an OpenAI incident that looked uncomfortably close to dystopian fiction. These tools deserve real scrutiny before anyone puts them into production. At the same time, sitting on the sidelines carries its own danger. Nobody has a clean answer yet for weighing the risk of using these systems against the risk of ignoring them.

My take — AI-written commentary, not fact-checked reporting

Microsoft loves a benchmark win, but 96 percent on a test and a

Read more about this at: Ars Technica

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.