Microsoft says its AI beat Anthropic's top security model. read the fine print first

the 96% score is real. the 'we beat Mythos' headline is doing a lot more work than the benchmark actually supports.

Aliteq
Priya Nair · Software & Systems Editor

Key point

Microsoft's MDASH harness, running its new MAI-Cyber-1-Flash model alongside OpenAI's GPT-5.4, scored 95.95% on the CyberGym vulnerability-hunting benchmark.

Key point

Anthropic's Mythos 5 scored 83.8% on the same benchmark — a 12-point gap Microsoft is calling a win.

Key point

The '50% cheaper' claim compares MDASH's new configuration to Microsoft's own previous best setup (GPT-5.4 + 5.4 mini + 5.3 codex) — not to Mythos's cost to run.

Key point

This sits inside Perception, Microsoft's new platform running 100+ AI agents in red, blue and green teams; a public preview is set for November 3, 2026.

Key point

CEO Mustafa Suleyman: 'We're shipping this into production immediately.'

The honest verdict

Real progress in automated vulnerability hunting, presented with more marketing spin than the underlying comparison can support. Worth watching Perception's public preview in November before drawing…

Aliteq

Read the full story

Microsoft says its AI beat Anthropic's top security model. read the fine print first

Read the full story on Aliteq