ALITEQ.

Microsoft says its AI beat Anthropic's top security model. read the fine print first

the 96% score is real. the 'we beat Mythos' headline is doing a lot more work than the benchmark actually supports.

Priya NairUpdated Aug 37 min readWeb story
security analyst monitoring code and alerts across multiple screens

Microsoft says its cybersecurity AI system just beat Anthropic's best security model by 12 points on the industry's toughest vulnerability-hunting benchmark, at half the cost. That's a real, verified number — 95.95% on CyberGym versus Mythos 5's 83.8%. It's also not the comparison the headline makes it sound like, and the fine print is the actual story.

What actually got measured

CyberGym tests whether an AI system can find and correctly triage real, previously-known vulnerabilities across large codebases — the kind of grinding work that usually eats a security team's week. Microsoft's MDASH harness paired its first purpose-built cybersecurity model, MAI-Cyber-1-Flash, with OpenAI's GPT-5.4 inside a multi-agent pipeline and posted a 95.95% score, first reported in detail by MarkTechPost. That's a strong result, and Microsoft's own announcement frames it as roughly 12 points ahead of Anthropic's Mythos 5.

95.95%

MDASH (MAI-Cyber-1-Flash + GPT-5.4)

CyberGym benchmark score

83.8%

Anthropic Mythos 5

CyberGym benchmark score

50%

Claimed cost reduction

vs. Microsoft's own prior MDASH config

Nov 3, 2026

Perception public preview

The catch: this isn't model versus model

Here's what the 12-point headline skips. MDASH isn't one model — it's Microsoft's own specialized cybersecurity model stacked with a second, unrelated foundation model from a different company, running inside a harness Microsoft built specifically for this benchmark. Mythos 5's 83.8% score, as far as the public materials show, is Anthropic's model operating without that kind of bespoke multi-model scaffolding. Comparing a purpose-built two-model system against a single model isn't dishonest, exactly, but it's not the fair fight the marketing chart implies either. Stack any capable model inside a harness tuned for one specific benchmark and you'd expect it to climb — that's what the harness is for.

CyberGym scores — same chart, different setups

MDASH (2 models + custom harness)95.95%

MAI-Cyber-1-Flash + GPT-5.4

Mythos 5 (single model)83.8%
lines of code on a monitor during a vulnerability scan
Perception's red, blue and green agent teams mirror how a real security operations center is structured. · Unsplash

Why this still matters, caveat and all

Strip away the head-to-head framing and there's a real trend underneath: multi-agent security tooling is starting to look production-viable, not just a demo. Perception's red-team-attacks, blue-team-defends, green-team-fixes structure mirrors how an actual security operations center runs, just with over 100 agents instead of a night shift of three analysts. In a year that's seen a genuinely brutal run of critical disclosures — Microsoft's own record 569-CVE Patch Tuesday in July, the SharePoint exploitation wave that hasn't really stopped — automated systems that can find and fix bugs at this hit rate, before disclosure rather than after, are worth taking seriously regardless of which vendor's chart you trust more.

My honest take: I'd trust the 12-point gap a lot more if Microsoft published Mythos's exact configuration alongside its own, so anyone could check whether the comparison is actually fair rather than taking Microsoft's word for the framing. There's a bigger tension sitting underneath this whole category, too — Anthropic itself recently disclosed that Claude models were used to breach real companies' systems during its own security testing. Every major AI lab is now racing to prove its models are best at finding and exploiting vulnerabilities, while simultaneously promising those same capabilities stay contained. Both things can be true. It's still worth sitting with.

We're shipping this into production immediately.

Mustafa Suleyman, Microsoft AI CEO

Is MAI-Cyber-1-Flash available to the public?
It's running inside Microsoft's own MDASH harness now; Perception, the broader multi-agent platform built on top of it, has a public preview scheduled for November 3, 2026.
What is CyberGym?
A benchmark that tests whether an AI system can find and correctly triage real, previously known software vulnerabilities in large codebases.
Did Microsoft's AI literally beat Anthropic's Mythos?
Its combined system scored 12 points higher on one benchmark. That's not the same as one model outperforming the other in a like-for-like test — Microsoft's score comes from a two-model harness built for the benchmark.
Is it really half the cost of Anthropic's model?
No — the 50% cost reduction Microsoft cites is compared to its own previous best MDASH configuration, not to what Mythos costs to run.
Should a security team switch tools based on this benchmark alone?
Not on this data alone. It's a promising internal result from one vendor's own testing, not an independent, apples-to-apples comparison.

Verdict

The honest verdict

Real progress in automated vulnerability hunting, presented with more marketing spin than the underlying comparison can support. Worth watching Perception's public preview in November before drawing any real conclusion about which vendor's approach actually wins.

Best for: Security teams evaluating AI-assisted vulnerability scanning tools

What's actually worth watching between now and November: whether Anthropic responds with its own CyberGym number under a disclosed, comparable setup, and whether Perception's public preview holds up once security researchers outside Microsoft get to poke at it with vulnerabilities the benchmark didn't already know about. Benchmark charts age fast in this industry — the same price war reshaping what AI costs to run is exactly the kind of pressure that produces a flattering chart before it produces a genuinely independent one.

Software & Systems Editor

Priya Nair

Priya has daily-driven more Linux distros than she can name and treats her setup like a workshop. She covers the operating systems, apps and settings worth your time — and cheerfully calls out the 'optimizations' that just quietly break your machine.

Work out the hardware

The Aliteq brief

The tech worth knowing — hardware, AI, gaming, deals. No spam, unsubscribe anytime.

Keep reading