Mustafa Suleyman stood in front of a small San Francisco crowd and said, with barely concealed satisfaction, that Microsoft’s new cybersecurity model beats Gemini, GPT 5.5 Cyber, GPT 5.6 Sol, and Anthropic’s Mythos 5 on Cyber Gym, the benchmark the industry actually uses. “The golden benchmark,” he called it. That’s not a subtle move. That’s a direct shot at two of Microsoft’s closest AI partners and competitors in the same breath.
According to TechCrunch, Microsoft on Monday launched MAI-Cyber-1-Flash, its first cybersecurity-specialized model, alongside a new agentic security platform called Perception. MAI-Cyber-1-Flash is built specifically to find vulnerabilities in complex codebases. It powers MDASH, Microsoft’s internal system for software vulnerability identification and remediation. The company claims the model is both more capable and more cost-effective than what competitors currently offer.
Perception is the broader platform play. It deploys teams of AI agents across three roles: red teams that simulate potential attacks and profile likely threat actors, blue teams that detect and triage existing bugs, and green teams that take corrective action against those bugs. Dave Weston, the lead engineer on Perception, described what that means in practice: work that used to take hours across multiple specialized security staff now produces a fix in minutes. Detection, prioritization, posture fixing, even a code fix. That’s a meaningful efficiency claim if it holds up in real enterprise environments.
The context matters here. Cybercriminals are increasingly using AI to accelerate and personalize attacks, and enterprise security teams are stretched. Hayete Gallot, Microsoft’s VP for security, framed Perception as a way to “defend against AI with AI at the scale and speed that the attackers have.” That framing is sharp because it’s accurate. The threat surface has grown faster than most security organizations can staff for.
But Microsoft is entering a field that’s already moving. Anthropic launched Mythos earlier this year through its Glasswing partner program, and OpenAI shipped its own security solution in May under a program called Daybreak. Both are targeting the same enterprise buyers. The difference with Microsoft is distribution. No company has deeper hooks into enterprise IT infrastructure, and pairing Perception with tools like Microsoft Defender or Azure gives it a natural adoption path that Anthropic and OpenAI simply don’t have.
The new tools are expected to be available in preview on November 3. The key questions between now and then are whether the benchmark performance translates to real-world detection rates, and whether Perception’s agentic workflows are reliable enough for security teams to trust at speed. In security, a false negative isn’t just a miss. It’s an incident.




