Copilot

Microsoft Opens Its MAI Audio and Image Models in Foundry

On April 2, 2026, Microsoft put three first-party MAI models into Foundry preview: half-cost transcription, a minute of speech in under a second, and a top-three image model.

Microsoft Opens Its MAI Audio and Image Models in Foundry — article cover
On this page6 SECTIONS
  1. Three Models Walk Into Foundry
  2. Transcription and Voice: A First-Party Audio Stack
  3. MAI-Image-2: A Top-Three Debut
  4. Pricing and the Responsible AI Fine Print
  5. What It Means for Developers and Product Teams
  6. Sources

On April 2, 2026, Microsoft AI CEO Mustafa Suleyman announced three first-party foundation models into public preview on Microsoft Foundry and MAI Playground: MAI-Transcribe-1 for speech-to-text, MAI-Voice-1 for voice generation, and MAI-Image-2 for image generation. The batch marks Microsoft’s in-house AI capabilities stepping beyond text into audio and vision — and it arrives positioned squarely as enterprise-grade APIs rather than consumer demos.

Suleyman’s framing is worth noting: these are models built from scratch for the enterprise, and the launch copy leans on the same weights powering both Microsoft’s own products and what developers rent by the hour. For teams already committed to the Azure stack, that collapses two vendor questions — capability and compliance — into one.

Three Models Walk Into Foundry

Microsoft’s Tech Community post frames Foundry as “the most complete AI and app agent factory,” with the three MAI models as the newest tiles. MAI-Transcribe-1 and MAI-Voice-1 deploy onto existing Azure AI Speech resources, so developers don’t rebuild their pipelines; MAI-Image-2 lives in Foundry’s standard model catalog. The same models also power Copilot, Bing Image Creator, PowerPoint, and Azure Speech — first-party products and external APIs share identical weights.

Transcription and Voice: A First-Party Audio Stack

MAI-Transcribe-1 leads with accuracy across 25 languages: first place in 11 of the top 25 languages on the FLEURS benchmark, beating Whisper-large-v3 in the remaining 14, and beating Gemini 3.1 Flash in 11 of those 14. Microsoft claims roughly 50% lower GPU cost than “leading alternatives,” batch transcription 2.5x faster than Azure Fast, and pricing from $0.36 per hour. It already handles transcription and dictation in Copilot’s Voice Mode — meaning the model has been absorbing production traffic at Copilot scale before external customers ever see it. Microsoft also acknowledges the model embeds documented human biases from training data and addresses them openly — a disclosure you rarely see in a launch post, and one that matters if your transcripts feed downstream automated decisions.

MAI-Voice-1 targets emotional range and speaker identity over long-form content; Microsoft describes writing speech with it as “meant to feel more like drawing than dictating.” A single GPU produces 60 seconds of audio in under one second, at $22 per million characters. The Personal Voice feature clones a voice from a 10-second sample, but only through an approval process — deliberate friction. Azure Speech’s existing gallery of 700+ voices remains available alongside it.

MAI-Image-2: A Top-Three Debut

MAI-Image-2 debuted at number three on the Arena.ai image leaderboard, with Microsoft calling the family firmly top-three and at least twice as fast as its predecessor. The copy leans hard on in-image text rendering — generating readable text inside images, historically a weak spot for diffusion models. WPP is an early enterprise adopter; global Chief Creative Officer Rob Reilly called it “a genuine game-changer” for campaign-ready work. Copilot, Bing, and PowerPoint will move onto it in phases, which means the consumer-facing image experience across Microsoft’s surfaces is about to converge on the same engine developers can call from Foundry.

Pricing and the Responsible AI Fine Print

The sticker prices: $0.36 per hour for transcription, $22 per million characters for voice, and $5 per million tokens of text input plus $33 per million tokens of image output for generation. Against prevailing API prices, Microsoft is clearly using first-party infrastructure to push unit costs down — the same front as the model-economics war we traced in our Replit free-mode analysis: whoever drives the marginal cost of intelligence lowest wins distribution. The fine print matters too: voice cloning requires approval, and the image and voice models ship with built-in guardrails. Responsible AI constraints are designed into the product, not bolted on afterward.

What It Means for Developers and Product Teams

Three practical effects. First, audio pipelines gain a credible supplier option — transcription and voice generation share one Azure Speech resource, which keeps migration costs low. Second, “the same model powers first-party products” means the usage scale is real: Copilot’s daily traffic is effectively a stress test for these weights. Third, the image API’s output pricing ($33 per million tokens) belongs in your selection spreadsheet, especially for marketing workloads that need heavy in-image typography — historically the failure mode that forced a second pass through an editing tool, and now a first-class capability of the model itself.

Sources

AI-assisted summary compiled from the sources above, reviewed by a human before publishing.

FOUND_THIS_USEFUL?

Support more practical AI articles, tutorials, and build notes.

BUY_ME_A_COFFEE
SHAREXEMAIL