Rohan Paul

@rohanpaul_ai

Low-resource ASR (automatic speech recognition) usually gets treated as a model problem. VoiceArena just launched Monsoon, and shows it's a data collection problem. Monsoon is a 50-language speech dataset built from unscripted conversations across 23 countries. Off the shelf, Whisper Medium scores 92.7% semantic WER on IndicVoices Telugu. VoiceArena fine-tuned the same 769M checkpoint on Monsoon and got 16.1%, narrowly ahead of MAI-Transcribe-2 and Gemini 3.1 Pro in its own tests. There was no architecture change and no bigger model. The 76.6-point drop comes entirely from what the model heard during fine-tuning. Most speech data pipelines filter noise out. Monsoon keeps it in on purpose. Short clips, noisy rooms and difficult acoustic conditions stay in the training distribution, and the mix isn't left to chance. VoiceArena stratifies segments by DNSMOS. That makes the share of degraded audio a controlled property of the data
打开原帖#511482
  1. Industry

    Alexandr Wang: S for superintelligence 😎
  2. Industry

    Alexandr Wang: seriously awesome use case for muse — clean up the neighborhood!
  3. Industry

    Yoshua Bengio: Automating more and more of the AI R&D could lead to AI progress radi…