AI

Microsoft MAI-Transcribe-2-Streaming hits $0.54 per audio hour

· Geeknewz Author

Silver condenser microphone in a recording studio

Microsoft AI shipped a voice-agent toolkit on October 1, 2026: MAI-Transcribe-2-Streaming for live speech-to-text, plus MAI-Voice-2.1 and MAI-Voice-2.1-Flash for text-to-speech. The pitch is simple. Hear the caller while they are still talking, start reasoning mid-sentence, and answer with a voice that can switch languages without swapping speaker identity.

That is more than a model card refresh. Microsoft is assembling the listen and speak ends of a voice loop on its own MAI stack, available through Microsoft Foundry, the MAI Playground, Vercel, Azure Voice Live, and (for the voice models) OpenRouter. Here is the price and latency math you can check, what the Artificial Analysis ranking actually claims, and who should care.

Black handheld microphone against a dark background
Photo via Unsplash (https://unsplash.com/photos/1510915228340-29c85a43dcfe). Free license.

Three models, three jobs

ModelJobSticker price (intro)Latency / scope claim
MAI-Transcribe-2-StreamingReal-time speech to text$0.54 per audio hour through year-end60 languages; first partials just over 100 ms of audio; ~320 ms first hypotheses in some coverage
MAI-Transcribe-2 (batch)Wait-until-done transcription$0.10 per audio hour (prior launch)Same language story; no continuous partials
MAI-Voice-2.1Expressive TTS$22 per 1M characters23 languages / 26 locales; same speaker across languages
MAI-Voice-2.1-FlashFast / cheaper TTS$15 per 1M charactersUp to 45 seconds of audio at ~150 ms end-to-end latency (Microsoft)

Our math on the transcription split: $0.54 ÷ $0.10 = 5.4× the batch model's intro rate. You are paying that premium for continuous partial transcripts over a WebSocket-style stream, not for a different language list. Flash TTS is $15 / $22 ≈ 68% of the full Voice-2.1 character price, which matches Microsoft's "about 60% cheaper than comparable models" marketing band if you treat Flash as the volume SKU.

What the leaderboard claim is

Microsoft says MAI-Transcribe-2-Streaming ranks first for both final and first-partial accuracy on Artificial Analysis's streaming speech-to-text board (chart dated September 28, 2026 in the company post). Unite.AI's write-up of that chart lists about a 2.5% final word-error rate, 2.8% first-partial rate, and 0.13 seconds to final transcription on the vendor-cited board. Artificial Analysis describes AA-WER Streaming as real-time chunked audio across roughly eight hours from AA-AgentTalk, VoxPopuli, and Earnings22. Those are vendor and third-party numbers, not Geeknewz lab results. Treat them as the public scoreboard Microsoft wants you to see, then verify on your own audio before you rip out a production ASR stack.

Microsoft Learn is clearer on the product shape: public preview, no SLA, intermediate results while the speaker talks, finals when a segment commits. Integration paths are an OpenAI Realtime-compatible WebSocket API and the Azure Speech SDK. SiliconANGLE separately notes the streaming SKU on Vercel AI Gateway at the same $0.54/hour intro price, and reminds readers that last month's non-streaming MAI-Transcribe-2 sat at $0.10/hour because it waits for the utterance to finish.

Closing the voice loop

A voice agent needs listen, decide, and speak. Streaming transcription covers listen. Voice-2.1 or Flash covers speak. The middle is still a reasoning model (Microsoft points at its own Mai-Thinking line in secondary coverage). Pairing Streaming with Flash is Microsoft's recommended latency sandwich: partials arrive early enough to start tool calls, and Flash answers in about 150 ms end-to-end for short clips so the conversation still feels human. Both voice models support cloning from a few seconds of reference audio with consent guardrails Microsoft says are built in. A 4,000-listener Turing-style test in the company materials put combined MAI-Voice ratings at 50.3% "equal or more human-like than human recordings," which is a marketing statistic, not a safety cert.

Microsoft also launched Chatter in the MAI Playground so you can hear the loop without wiring Foundry first. That is the right place to start if you are evaluating UX, not unit economics.

Geeknewz verdict

Geeknewz's view: if you are building call-center agents, live captions, or multilingual tutors and you already sit in Azure or Foundry, try Streaming + Flash on a narrow pilot before year-end intro pricing ends. The 5.4× jump over batch transcription is only worth it when partials actually unlock mid-utterance actions. If you only need overnight meeting notes, stay on the $0.10/hour batch model. And keep the preview warning in mind. Learn marks Streaming as public preview without an SLA, so do not bet a regulated contact center on it until Microsoft graduates the SKU and you have measured error rates on your accents and domains.

Source: Microsoft Learn; company launch details via Microsoft Source and Unite.AI; additional reporting from SiliconANGLE.