Microsoft AI released two new in-house models into public preview on Wednesday — MAI-Image-2.5-Pro, its highest-fidelity image generator to date, and MAI-Voice-2-Flash, a speech model built for high-volume enterprise workloads. The announcement from Microsoft AI's Superintelligence team signals a major shift: the company is moving away from leaning on OpenAI's frontier models for routine tasks. By deploying these homegrown tools across Bing, PowerPoint, OneDrive, Dynamics 365, Excel, GitHub Copilot, and Azure, Microsoft is proving it can power its own products with independent, cost-efficient infrastructure.
Two models for two very different needs
The company is positioning these releases at opposite ends of the quality-speed-cost curve. MAI-Image-2.5-Pro targets the premium tier, focusing on hero imagery and precise in-image text rendering—a known weakness for many generative models. It is priced at $5 per million text input tokens, $8 per million image input tokens, and $106 per million image output tokens. Meanwhile, MAI-Voice-2-Flash is built for the high-volume, low-latency needs of call centers and voice agents. It runs twice as fast as MAI-Voice-2 and costs 32% less, priced at $15 per million characters.
Production data shows massive cost savings
The most compelling reason for enterprise buyers to consider these models is the reported reduction in overhead. Microsoft shared internal production metrics showing that swapping out third-party frontier models for these in-house options significantly lowers expenses:
- PowerPoint: MAI-Image-2.5-Pro reduces GPU costs by up to 84% compared with OpenAI's GPT-Image-2.
- Dynamics 365 Contact Center: MAI-Voice-2-Flash reduces GPU costs by up to 89%.
- OneDrive: Switching to the MAI-Image-2.5 family resulted in a 26% increase in save rates, 25% lower P95 latency, and 2.5x greater efficiency under medium-utilization workloads.
Better performance in critical workflows
These models aren't just cheaper; they are performing in high-stakes environments. In healthcare, Dragon Copilot—which handles 28 million patient encounters last quarter for 170,000 medical providers—now uses MAI-Transiap-1.5 for multilingual workflows across 58 languages. Microsoft reports a 50% relative reduction in transcription and language-identification error rates across most languages. For developers, MAI-Code-1-Flash in GitHub Copilot achieves an 10% higher code accept rate than GPT-5.4 Mini and Claude Haiku 4.5 while using 10% fewer median tokens.
The strategy: move away from 'frontier' for routine tasks
CEO Satya Nadella framed this as a move toward "Frontier Diffusion & Control." The goal is to use frontier models for complex needs while routing routine traffic to optimized internal models. By training models like MAI-Code-1-Flash in Excel reinforcement learning environments, Microsoft has created a tool on par with GPT-5.6 for common spreadsheet tasks that can run on older Nvidia H100 and A100 GPUs. This allows the company to save its newest GB200 clusters for training rather than serving. For your business, this means high-quality AI becomes more accessible and less expensive to run as Microsoft absorbs the heavy lifting of routine AI infrastructure.
If you are currently evaluating AI costs for your enterprise, start by auditing your high-volume, repetitive tasks—like call center scripts or spreadsheet reformatting. These are the areas where switching to Microsoft';s in-house models will provide the most immediate budget relief.

