Rohan Paul
@rohanpaul_ai
Low-resource ASR (automatic speech recognition) usually gets treated as a model problem.
VoiceArena just launched Monsoon, and shows it's a data collection problem.
Monsoon is a 50-language speech dataset built from unscripted conversations across 23 countries.
Off the shelf, Whisper Medium scores 92.7% semantic WER on IndicVoices Telugu. VoiceArena fine-tuned the same 769M checkpoint on Monsoon and got 16.1%, narrowly ahead of MAI-Transcribe-2 and Gemini 3.1 Pro in its own tests.
There was no architecture change and no bigger model. The 76.6-point drop comes entirely from what the model heard during fine-tuning.
Most speech data pipelines filter noise out. Monsoon keeps it in on purpose.
Short clips, noisy rooms and difficult acoustic conditions stay in the training distribution, and the mix isn't left to chance. VoiceArena stratifies segments by DNSMOS. That makes the share of degraded audio a controlled property of the data