Meta Superintelligence Labs
@alexandr_wang
today we're rolling out muse voice transcribe, our first real-time audio perception model - SOTA in streaming speech-to-text. also handles speaker diarization and endpointing natively in a single model.
Muse Voice Transcribe is an autoregressive multimodal model from the Muse Spark family. It delivers real-time streaming ASR, diarization with 20+ speakers, and endpointing. It is multilingual with seamless code-switching and improves accuracy with language, keyword, and context biasing. Meta ranks first on Artificial Analysis on streaming speech-to-text and on public diarization benchmarks as of September 1, 2026.
It is available today via Meta Model API, Meta AI for Mac, and Muse Code.