Nvidia released Nemotron 3 Diarization. On September 27, The Decoder reported that the model decides who is speaking at a given moment. It has about 100 million parameters, and the weights are free to take. It can separate up to eight speakers and notice when several people talk at once. More people, heavy background noise, or reverberation push the error rate up.

Paired with a speech recognizer such as Parakeet, it can label a transcript with speakers. The labels are anonymous, such as speaker_2. It works on recordings and on live audio. This piece did not open the leaderboard. The figures come from The Decoder, which points to Hugging Face.

[1]
Copperplate of eight empty chairs, with a darker shadow on only one.
Eight chairs, and only one carries a heavier shadow. It stands for the person speaking among up to eight, and it is not a photograph of a meeting., AI-generated illustration, not a news photograph

On VoiceArena's Diarization-Bench, the report says this free model is currently first, with a 14.72 percent error rate, ahead of the next system at 19.3 percent. The bench counts overlapping speech, and it scores even small timing misses at speaker changes as errors. Against the previous Streaming Sortformer, and with a 1.04-second buffer, the error rate falls by 41 percent on average across eight test scenarios. The same article also writes a DER of 14.7 percent and a 41 percent lead on that bench. Both wordings stay: 14.72 percent is the listed error rate, and 41 percent is the average drop on eight scenarios at a 1.04-second buffer, not a drop in every room.

The audio buffer has four settings, from 30.4 seconds down to 0.32 seconds. Shorter buffers generally lose accuracy.

[1]

The figures do not say whether the audio was a meeting, a phone call, or a studio, and they do not say the model is useless past eight speakers. The report already states the boundary: more people, noise, and reverberation raise errors, and a shorter buffer generally lowers accuracy. First place on that bench is not a claim that every room will come out labeled correctly. Free weights also do not mean the model is already wired into a particular transcription product. This piece did not run its own test.

[1]

要点

  • Nemotron 3 Diarization has about 100 million parameters, free weights, and separates up to eight speakers.
  • VoiceArena lists a 14.72 percent error rate, next best 19.3 percent. At a 1.04-second buffer, eight scenarios average a 41 percent drop.
  • Buffers run from 30.4 seconds to 0.32 seconds. Shorter ones generally lose accuracy.
  • More people, noise, or reverberation raise the error rate. Speaker labels are anonymous.