Microsoft MAI Models: Voice Agents with 100ms Latency

Alex da Cruz
Alex da Cruz is a full-stack developer based in São Paulo, Brazil. He works with React, TypeScript and automation, and uses AI daily to solve real problems in code and operations — not as a demo. He has run an e-commerce operation end to end, and now builds and maintains the automation pipeline behind this blog. He writes about what he actually tests.
According to Microsoft, the new MAI-Transcribe-2-Streaming model delivers partial results in just over 100 milliseconds across 60 languages. This speed allows voice agents to process user speech and respond while the user is still mid-sentence, removing the awkward pauses common in automated phone trees.
How much do the new Microsoft voice models cost?
Through the end of the year, transcription costs $0.54 per hour of audio under an introductory price. For text-to-speech, the MAI-Voice-2.1-Flash variant achieves a latency of 150 milliseconds at a rate of $15 per million characters, down from the standard $22 tier.
What changes for customer service automation?
Both voice models can clone a target voice using only a few seconds of reference audio, while integrated safeguards attempt to prevent unauthorized use. In tests cited by Microsoft, roughly half of 4,000 participants could not distinguish the generated voice from a real human speaker. Developers can access these models immediately through Microsoft Foundry, the MAI Playground, and OpenRouter for integration into production workflows.
Sources
Frequently asked questions
- What is the latency of MAI-Transcribe-2-Streaming?
- The model delivers its first partial results in just over 100 milliseconds, allowing voice agents to respond while users are still speaking.
- How much does the new Microsoft transcription cost?
- Through the end of the year, transcription costs $0.54 per hour of audio under the introductory pricing.
- Where can developers access MAI-Voice models?
- The models are available through Microsoft Foundry, the MAI Playground, and OpenRouter.
Comments
0 comments
Be the first to comment.
Continue Lendo

Manus 2.0: mobile control, video editing and AI agents
Manus 2.0 adds mobile remote control, video editing, multiplayer game hosting, and cuts agent running costs by 32 percent.

Shopify WebMCP Checkout: AI Agents Complete Purchases
Shopify's new WebMCP checkout support allows browser-based AI agents to complete purchases natively using structured APIs.

Nvidia SoL-Pi: How to Cut Coding Agent Token Costs in Half
Nvidia's SoL-Pi system optimizes coding agent control layers, cutting token consumption by up to 49% and lowering hourly execution costs.