A night-shift call-center agent with a headset watches multilingual captions scroll on a monitor while a wall placard lists $0.10 per hour
Transcription stops looking like a purchase-order fight and starts looking like a meter., AI-generated illustration, not a news photograph

On September 3, Microsoft AI shipped MAI-Transcribe-2 with a headline that leaves little oxygen: it calls the model the world’s fastest, most accurate, and cheapest speech recognizer. The hard numbers in the post: 5.2% average WER atop FLEURS across 60 languages; the accuracy–latency Pareto frontier on Artificial Analysis, second on that board’s WER ranking; roughly 10× / 7× / 5× batch throughput versus GPT-Transcribe, ElevenLabs Scribe v2, and Gemini 3.5 Transcribe; and a launch price of $0.10 per audio hour, explicitly a limited-time offer through year-end.

TradePoint lays out the price history: the line opened near $0.36 five months ago, so this cut is about 72%. For 100,000 call-center hours a year, the illustrative bill drops from $36,000 to $10,000—nobody argues about transcription in a procurement meeting after that.

[1][2]

Features that used to be upsells

The capability list reads like a specialty vendor’s add-on menu emptied into the base rate: diarization, word-level timestamps, keyword biasing, verbatim/clean styles, code-switching pairs such as Hinglish and Spanglish, automatic language ID, noisy-room robustness, and 60 languages. TradePoint clocks the cadence—25 languages in April, 43 in June, 60 in September—three releases in five months, each welding paid extras into the floor.

Availability sits on Microsoft Foundry, MAI Playground, and Open Router. Named rivals are frontier-lab ASR: GPT-Transcribe, Gemini 3.5 Transcribe, Whisper V3-Large, Scribe v2. Deepgram, AssemblyAI, Speechmatics, and Rev never appear in the launch copy—yet a dime-an-hour floor will still hit their high-volume contracts.

[1][2]

How to read the scoreboard

FLEURS averages need coverage context: 1.5 once posted about 3.7% average WER; 2.0’s 5.2% more likely reflects a longer language table and low-resource drag than a pure regression—buyers should demand per-language sheets. Artificial Analysis’s independent API numbers are the second pillar Microsoft leans on: climbing from third to second on WER reportedly means clearing ElevenLabs; owning the Pareto frontier means no rival is both more accurate and faster.

What the big type leaves soft: public materials emphasize batch throughput more than real-time streaming parity, and $0.10 is promotional. The permanent price is still a blank cell.

[1][2]

Paying less rent to partners

TradePoint frames MAI inside Microsoft’s wider move: after repeated revisions to its OpenAI terms, the company is using in-house models to cut model rent. Transcription sits on native pipes—Teams, Nuance clinical notes, Azure Speech—so this is less a demo toy than a way to pull captive traffic off someone else’s price list.

[2]

Take

I read this as a floor-price war in ASR, not another SOTA parade. If the promo snaps back toward $0.36 after December, the story shrinks to a marketing pulse. If the floor holds, specialty vendors either sell diarization-plus-compliance workflows as the real product or retreat into verticals Microsoft never named.

Buyers should stop retweeting “cheapest on earth” and run their own noisy call-center samples for per-language WER—then write the promo end date into the contract as a hard clause.

[1][2]