The Decoder reported on October 2 that Microsoft released two text-to-speech models along with a transcription model. MAI-Voice-2.1 is meant to speak 23 languages in the same voice, with a native accent in each language. Microsoft says the MAI-Voice-2.1-Flash variant has a latency of 150 milliseconds and costs $15 per million characters instead of $22. Those figures are Microsoft’s. The article does not include an independent measurement.

[1]
A blank circle with no face, above a column of short strokes that are dark on the left and pale on the right.
The circle has no face. Each short stroke below it is dark on the left and pale on the right. That is a test in which about half the listeners heard a real person. An illustration, not a test screenshot., AI-generated illustration, not a news photograph

Both voice models can clone a voice from just a few seconds of reference audio. The article says built-in safeguards are meant to prevent misuse. This piece keeps those two statements and does not describe how the audio is supplied or how a safeguard could be bypassed. The models are available through Microsoft Foundry and the MAI Playground, and the two voice models are also on OpenRouter. In one test, about half of 4,000 participants thought the voices belonged to a real person. The article does not say which model was tested, which language was used, or what task the participants heard.

[1]

The same article also describes the real-time transcription model MAI-Transcribe-2-Streaming. Microsoft says it ranks first for accuracy on Artificial Analysis, transcribes 60 languages, and returns its first partial results in just over 100 milliseconds. Microsoft says this lets a voice agent respond while the other person is still mid-sentence. Through the end of the year, the introductory price is $0.54 per hour of audio. The first-place claim, the language count, and the 100 milliseconds are Microsoft’s figures. This is The Decoder’s account of the release, not Microsoft’s own blog post.

[1]

要点

  • MAI-Voice-2.1 speaks 23 languages in one voice. Flash latency is 150 milliseconds, at $15 per million characters instead of $22.
  • Both voice models can clone from a few seconds of reference audio and include safeguards against misuse. The method is not described.
  • About half of 4,000 participants thought the voices were a real person. The article does not name which model was tested.
  • The same article’s streaming transcription model, on Microsoft’s figures, covers 60 languages, returns a first partial in just over 100 milliseconds, and costs $0.54 an hour through year end.