MAI-Cyber-1-Flash was built to process as much as 90% of MDASH workloads. The remaining 10%, which includes the most difficult tasks, is escalated to GPT-5.4. Microsoft AI CEO Mustafa Suleyman said the advantage comes from the orchestration system rather than any one model.
“The harness is like a router,” Suleyman told VentureBeat. “It’s kind of like guardrails and a rule set of an organizing logic, which matches queries to... incoming problems to a model that suits the problem.”
He described the setup as three connected pieces: the routing layer, MAI-Cyber-1-Flash for faster and less expensive work, and GPT-5.4 for more demanding coding problems. The broader MDASH system manages long-running workflows that can involve stored context, external databases, code generation, validation, and repeated handoffs between models.
“These are very complicated, long, agentic loops which require storing state, drawing on another database, consulting best practice... handing back to a small model, writing a bunch of code, validating that that was correct,” Suleyman said. “There’s like hundreds of steps to solve that, and that’s why it’s really the system together that delivers the better performance.”
Cost was a central factor in Microsoft’s model selection. Suleyman said GPT-5.4 offered a stronger balance of capability and price than more expensive frontier systems. “GPT-5.6 is expensive. GPT-5.4 is incredibly good relative to its cost,” he said. “The whole game here is to reduce the costs. Mythos and so on are extremely expensive models... we want to be able to deliver better performance for cheaper. That’s what customers want.”
Project Perception extends that system-level approach into coordinated security operations. It uses red-team agents to search for attack paths, blue-team agents to investigate and prioritize risks, and green-team agents to repair vulnerabilities and strengthen defenses.
Microsoft is positioning the launch around a broader shift in enterprise AI spending. Suleyman said companies initially relied heavily on the most capable models but later faced rising token expenses as usage expanded across their operations. He argued that chip availability and the cost of producing model output are becoming practical limits on adoption.
“The key barrier to adoption is access to chips, and cost is a function of chips,” he said. “No matter how much money you’ve got, there’s actually a limited supply of chips. Then trying to squeeze more model output on fewer chips is clearly super valuable.”
Microsoft also pointed to its security telemetry as an important part of the system’s development. The company processes more than 100 trillion security signals each day and draws information from 1.6 million customers. Suleyman described that history of security data and operational experience as a competitive advantage.
“We have trillions and trillions of data points going back decades,” he said. “It is, I think, the largest longitudinal cybersecurity dataset around.”
“That is definitely a moat for us,” Suleyman added. “Both the data and the expertise, and just the experience in the institution of going through that process.”
The benchmark results still carry limitations. Microsoft conducted the evaluation itself, and the comparison measured its complete agentic system against other model configurations rather than testing individual models under identical conditions. The result therefore reflects the performance of Microsoft’s routing, tools, and model combination, not MAI-Cyber-1-Flash alone.
Microsoft is also restricting access because a system capable of discovering software vulnerabilities could be useful to attackers. Suleyman said approved users must demonstrate legitimate intent and technical ability, while Microsoft will monitor API activity and expand access gradually.
“We’re very strict about who gets access to the model, and we’re very careful about that,” he said. “We constantly monitor the API and usage.”
The company said its AI Red Team evaluated the model through automated and expert-led testing, while an outside party conducted an independent assessment. Microsoft is also using tenant isolation, auditing, sandboxed execution, and environments without internet access.
Suleyman said Microsoft expects to continue building specialized models and combining them through shared orchestration systems. He described future work that could bring coding, voice, image, and transcription models into the same framework.
He also questioned whether enterprise AI will ultimately depend on a single large multimodal model. “It remains to be seen whether one giant model that is fully multimodal is actually able to deliver additional transfer learning benefit because of the integration,” he said, “or whether it’s just a big lumbering expensive giant.”
The new cybersecurity products reflect Microsoft’s alternative approach: use smaller models for routine work, reserve frontier systems for harder cases, and rely on orchestration and proprietary data to improve performance. In security, where Microsoft controls both large volumes of telemetry and widely used enterprise products, the company is betting that the surrounding system will matter more than the size of any individual model.
This analysis is based on reporting from Venture Beat.
Image courtesy of EPA Photo.
This article was generated with AI assistance and reviewed for accuracy and quality.