
Candidate headlines
- A dime an hour: Microsoft’s MAI-Transcribe-2 claims faster, cheaper, stronger ASR
- Third drop in five months: MAI-Transcribe-2 hits $0.10/hr and FLEURS #1
- Taking aim at OpenAI, Google, and ElevenLabs: Microsoft keeps cutting the rent on speech
Lead
On September 3, 2026, Microsoft AI launched MAI-Transcribe-2. The company calls it its most capable transcription model yet—and the most capable and efficient among competitors—while VentureBeat frames the release as undercutting OpenAI, Google, and ElevenLabs on price and speed.
Headline numbers in the launch materials: #1 on FLEURS across 60 languages at 5.2% average word-error rate (WER); defines the Artificial Analysis accuracy–latency Pareto frontier and ranks second on that firm’s WER leaderboard; batch speed cited as about 10× / 7× / 5× faster than GPT-Transcribe / Scribe v2 / Gemini 3.5 Transcribe; launch price $0.10 per hour of audio (Microsoft says a limited-time offer through year-end). Demos are live on Microsoft Foundry, MAI Playground, and Open Router.
Source brief
What Microsoft AI’s post says
Per Microsoft AI (2026-09-03):
- Positioning: claims the fastest, most accurate, and cheapest speech-recognition model; beats leading models such as Gemini 3.5 Transcribe, GPT-Transcribe, Whisper V3-Large, and ScribeV2 while handling a broader range of real-world audio.
- Benchmarks: #1 on FLEURS across 60 languages at 5.2% average WER; defines Artificial Analysis’s accuracy–latency Pareto frontier; ranks second on the Artificial Analysis WER leaderboard.
- Speed (Artificial Analysis evals, as cited): about 10× faster than OpenAI’s GPT-Transcribe, 7× vs ElevenLabs’ Scribe v2, and 5× vs Gemini 3.5 Transcribe, with higher accuracy; substantially lower latency especially on long-form audio, with up to ~10× faster processing than leading competitors.
- Price: $0.10 per hour, limited-time through year-end.
- Feature pack: speaker diarization, word-level timestamps, keyword biasing, configurable styles (verbatim keeps fillers/false starts; clean strips fillers), code switching (e.g., Hinglish, Spanglish), automatic language ID, noise robustness, 60 languages.
- Availability: Microsoft Foundry, MAI Playground, Open Router.
What VentureBeat adds on product and rivalry
VentureBeat (2026-09-03, 7:00 a.m. PT) adds:
- Price history: the line’s first model ~five months earlier was $0.36/hr; the new early-bird rate cuts that by roughly 72%. Example: 100,000 call-center hours/year could drop from about $36,000 to $10,000 at list rates.
- Language growth: 60 languages (43 in June’s 1.5; 25 in April’s original).
- How to read FLEURS: 1.5 reported 3.7% average WER; 2.0’s 5.2% likely reflects broader low-resource coverage rather than pure regression—buyers should request per-language breakdowns.
- Artificial Analysis: independent API testing; in June, 1.5 ranked third at ~2.4% WER behind Alibaba’s Fun-Realtime-ASR-preview and ElevenLabs’ Scribe v2; climbing to second suggests clearing ElevenLabs. “Pareto frontier” means no rival is both more accurate and faster.
- Cadence: three releases in five months (Apr 2 → Jun 2 → Sep 3), each expanding languages ~40% while folding premium-tier features into the base rate.
- Strategy frame: after repeated partnership amendments with OpenAI, in-house MAI models are read as reducing rent on external models; transcription is a natural swap target for Teams, Nuance clinical documentation, and Azure speech traffic. The piece also cites prior Mustafa Suleyman comments on a small focused team and GPU-cost advantages.
- Rivals not claimed beaten: accuracy narrative does not claim to top Alibaba-family numbers on independent boards; specialist vendors (Deepgram, AssemblyAI, Speechmatics, Rev) are not named head-on, but a $0.10 base rate pressures high-volume contracts.
Technical and product value
Why it matters (author judgment): Transcription procurement has long been unbundled—base rate plus add-ons for diarization, timestamps, and domain vocabularies. Folding those into a $0.10 base rate uses inference efficiency to subsidize feature bundling. For high-hour pipelines (contact centers, meetings, captions, compliance audio), unit price beats another leaderboard tick.
Implications for builders (author judgment):
- One multilingual SKU: 60 languages plus auto language ID and code switching reduce “one vendor per language” ops.
- Style toggles are compliance toggles: verbatim for legal/QA; clean for readable notes—same API, two SLAs.
- Keyword biasing aims at specialist moats: drug names, employee IDs, SKUs are exactly where vertical ASRs charge—still must test on your lexicon.
- Channel convenience ≠ enterprise closure: Foundry / Playground / Open Router ease trials; residency, retention, and training-use terms remain thin in the launch posts.
Competition and strategy
Named rivals (fact boundary): Microsoft names GPT-Transcribe, Gemini 3.5 Transcribe, Whisper V3-Large, and ScribeV2. VentureBeat notes specialist vendors are not directly named, and Microsoft does not claim to beat Alibaba-family accuracy on independent boards.
Strategic reading (author judgment):
- Modality-by-modality in-house stack: not one do-everything model, but specialized models optimized for inference cost, sold via Foundry and swapped into Microsoft products.
- Transcription is the cleanest rent-substitution proof: bounded problem, objective metrics, and captive audio from Teams / Nuance / Azure Speech—every migrated hour is an hour not paid to an external partner.
- Price war commoditizes ASR: five months from $0.36 to $0.10 with features bundled shifts differentiation toward domain depth and integrations, not list price alone.
- OpenAI relationship framing: VentureBeat reads the launch as evidence Microsoft would rather own the factory than rent the output after partnership amendments—that is press framing, not a contract excerpt.
Risks, limits, and controversies
- Vendor-cited third-party speed/accuracy: 10×/7×/5× and Pareto claims come from Artificial Analysis evals as relayed by Microsoft—check test mix (English business-heavy), batch vs streaming, and default API settings.
- Higher FLEURS average: vs 1.5’s 3.7% (43 langs), 2.0’s 5.2% (60 langs) may be diluted by low-resource languages; without per-language tables, do not treat it as uniformly “more accurate.”
- Streaming gap: launch materials stress batch throughput and long-form; VentureBeat flags silence on real-time ASR needed for voice agents and live captions.
- No diarization error metric: WER does not measure speaker attribution; meetings can be lexically right and speaker-wrong.
- Promo vs list price: Microsoft says the dime rate runs through year-end—get post-promo pricing in writing.
- Data/compliance unknowns: residency, retention, and training use are not detailed in the launch posts; regulated buyers need Foundry contract answers.
- Slogan-grade “fastest/most accurate/cheapest”: marketing umbrella; independent boards still show competitors (e.g., Alibaba-family accuracy) Microsoft does not claim to beat.
Critic’s take
My read: MAI-Transcribe-2 is less another speech-leaderboard poster and more Microsoft using inference efficiency → list price → full feature bundle to turn transcription from a margin center into infrastructure—save rent on captive audio first, then sell surplus capacity at a dime an hour.
Watch three production questions, not the “10×” slogan: your per-language WER, whether diarization holds, and the real bill after $0.10 expires. If those clear, specialist API pricing keeps getting compressed; if not, this remains a polished Foundry launch note.
Conclusion and 6–12 month outlook
Expect three tracks:
- Bundled ASR at dime-scale list prices becomes the default battlefield, forcing specialists and frontier labs to rewrite rate cards.
- Batch Pareto ≠ streaming win; real-time agents and live captions open a separate race.
- Microsoft’s modality playbook continues internal substitution: if transcription works, the same “specialized in-house model + Foundry + product swap” template spreads to more workloads.
Bottom line: MAI-Transcribe-2 pushes speech transcription toward commoditization with a FLEURS #1 narrative, Artificial Analysis speed claims, and a limited-time dime rate—buyers should decide on per-language, per-scenario, batch-vs-streaming tests, not launch adjectives.