Microsoft says its cybersecurity AI system just beat Anthropic's best security model by 12 points on the industry's toughest vulnerability-hunting benchmark, at half the cost. That's a real, verified number — 95.95% on CyberGym versus Mythos 5's 83.8%. It's also not the comparison the headline makes it sound like, and the fine print is the actual story.
What actually got measured
CyberGym tests whether an AI system can find and correctly triage real, previously-known vulnerabilities across large codebases — the kind of grinding work that usually eats a security team's week. Microsoft's MDASH harness paired its first purpose-built cybersecurity model, MAI-Cyber-1-Flash, with OpenAI's GPT-5.4 inside a multi-agent pipeline and posted a 95.95% score, first reported in detail by MarkTechPost. That's a strong result, and Microsoft's own announcement frames it as roughly 12 points ahead of Anthropic's Mythos 5.
95.95%
MDASH (MAI-Cyber-1-Flash + GPT-5.4)
CyberGym benchmark score
83.8%
Anthropic Mythos 5
CyberGym benchmark score
50%
Claimed cost reduction
vs. Microsoft's own prior MDASH config
Nov 3, 2026
Perception public preview
The catch: this isn't model versus model
Here's what the 12-point headline skips. MDASH isn't one model — it's Microsoft's own specialized cybersecurity model stacked with a second, unrelated foundation model from a different company, running inside a harness Microsoft built specifically for this benchmark. Mythos 5's 83.8% score, as far as the public materials show, is Anthropic's model operating without that kind of bespoke multi-model scaffolding. Comparing a purpose-built two-model system against a single model isn't dishonest, exactly, but it's not the fair fight the marketing chart implies either. Stack any capable model inside a harness tuned for one specific benchmark and you'd expect it to climb — that's what the harness is for.
CyberGym scores — same chart, different setups
MDASH (2 models + custom harness)95.95%
MAI-Cyber-1-Flash + GPT-5.4
Mythos 5 (single model)83.8%
Perception's red, blue and green agent teams mirror how a real security operations center is structured. · Unsplash
My honest take: I'd trust the 12-point gap a lot more if Microsoft published Mythos's exact configuration alongside its own, so anyone could check whether the comparison is actually fair rather than taking Microsoft's word for the framing. There's a bigger tension sitting underneath this whole category, too — Anthropic itself recently disclosed that Claude models were used to breach real companies' systems during its own security testing. Every major AI lab is now racing to prove its models are best at finding and exploiting vulnerabilities, while simultaneously promising those same capabilities stay contained. Both things can be true. It's still worth sitting with.
We're shipping this into production immediately.
Mustafa Suleyman, Microsoft AI CEO
Is MAI-Cyber-1-Flash available to the public?
It's running inside Microsoft's own MDASH harness now; Perception, the broader multi-agent platform built on top of it, has a public preview scheduled for November 3, 2026.
What is CyberGym?
A benchmark that tests whether an AI system can find and correctly triage real, previously known software vulnerabilities in large codebases.
Did Microsoft's AI literally beat Anthropic's Mythos?
Its combined system scored 12 points higher on one benchmark. That's not the same as one model outperforming the other in a like-for-like test — Microsoft's score comes from a two-model harness built for the benchmark.
Is it really half the cost of Anthropic's model?
No — the 50% cost reduction Microsoft cites is compared to its own previous best MDASH configuration, not to what Mythos costs to run.
Should a security team switch tools based on this benchmark alone?
Not on this data alone. It's a promising internal result from one vendor's own testing, not an independent, apples-to-apples comparison.
Verdict
The honest verdict
Real progress in automated vulnerability hunting, presented with more marketing spin than the underlying comparison can support. Worth watching Perception's public preview in November before drawing any real conclusion about which vendor's approach actually wins.
Best for: Security teams evaluating AI-assisted vulnerability scanning tools
What's actually worth watching between now and November: whether Anthropic responds with its own CyberGym number under a disclosed, comparable setup, and whether Perception's public preview holds up once security researchers outside Microsoft get to poke at it with vulnerabilities the benchmark didn't already know about. Benchmark charts age fast in this industry — the same price war reshaping what AI costs to run is exactly the kind of pressure that produces a flattering chart before it produces a genuinely independent one.