Microsoft priced MAI-Image-2.5-Pro at $5 per million text input tokens, $8 per million image input tokens and $106 per million image output tokens. The company described it as its most capable image model so far.
“MAI-Image-2.5-Pro is a strong leap forward for GenMedia tools,” said Rob Reilly, WPP’s global chief creative officer. “Beyond the impressive image quality, its ability to render text with this kind of accuracy is a real breakthrough. It also understands natural language edits, so creative iteration becomes faster and far more intuitive. Microsoft has firmly established itself among the leaders in generative AI.”
MAI-Voice-2-Flash, which Microsoft first introduced at Build, is twice as fast as MAI-Voice-2 and costs 32% less. It is priced at $15 per million characters while retaining the natural speech patterns and acoustic quality of the earlier model.
The launches are part of Microsoft’s broader effort to develop model families that can be matched to specific products and workloads. Rather than rely on one general-purpose system, the company is offering different combinations of quality, latency and operating cost.
Microsoft says those internally developed models are already running across Bing, PowerPoint, OneDrive, Dynamics 365 and Azure. The company began the effort with a focus on models trained using traceable, enterprise-grade data and without distillation from outside systems.
Bing Image Creator now runs entirely on MAI-Image-2.5, which has become the default model for both generation and editing. Microsoft says the model gives users more control over changes while keeping the image workflow fully in-house.
PowerPoint is also using MAI-Image-2.5 for image-to-image features. According to Microsoft, the deployment cuts GPU costs by as much as 84% compared with GPT-Image-2.
In OneDrive, MAI-Image-2.5 now handles several core image-editing tasks. Microsoft reports that the change raised save rates by 26%, lowered P95 latency by about 25% and produced 2.5 times greater efficiency under medium-utilization workloads.
The voice model is already running inside Dynamics 365 Contact Center, which customers including T-Mobile and EasyJet use to build call center agents. Microsoft says MAI-Voice-2-Flash can reduce GPU costs by up to 89% in that environment.
The model has also been added to Azure Voice Live, giving developers access to low-latency, expressive voices for speech-to-speech agents.
Microsoft is extending its in-house model strategy into healthcare as well. MAI-Transcribe-1.5 now supports multilingual workflows in Dragon Copilot, which is used by 170,000 medical providers and processed 28 million patient encounters in the previous quarter.
The transcription model supports 58 languages. In Microsoft’s internal testing, it reduced transcription and language-identification error rates by 50% across most languages compared with the model it replaced. The company also said early research showed better downstream accuracy in medical notes.
Together, the deployments show Microsoft moving its MAI models from research into large-scale product use. The company is emphasizing not only benchmark performance, but also lower serving costs, faster response times and measurable changes in how users interact with its products.
Both MAI-Image-2.5-Pro and MAI-Voice-2-Flash are available through Microsoft Foundry and the MAI Playground during public preview.
This analysis is based on reporting from Microsoft AI.
Image courtesy of Microsoft AI.
This article was generated with AI assistance and reviewed for accuracy and quality.