NVIDIA releases the free Nemotron-3 Diarization speech-segmentation model, with about 100 million parameters. It supports real-time and recording audio processing for up to 8 people speaking simultaneously. In the VoiceArena speech-segmentation benchmark test, the model topped the leaderboard with an error rate of 14.72%, significantly outperforming the system in second place. Compared with the previous Streaming Sortformer, the new model substantially reduces the average error rate by 41% with an audio buffer of 1.04 seconds, and supports configurable dynamic audio buffers ranging from 0.32 seconds to 30.4 seconds. The model can work with speech recognition systems such as Parakeet to automatically generate text with anonymous speaker labels. With this lightweight, high-precision speech-segmentation model, NVIDIA will provide underlying support for applications such as real-time meeting transcription and intelligent customer service, accelerating the deployment of multi-speaker recognition technology on edge devices and in real-time commercial scenarios.

Article author and source: AIBase

Nemotron-3-Diarization model release

On September 27, NVIDIA officially released a free AI model, Nemotron-3-Diarization, with about 100 million parameters. The model is focused on the speech diarization (Diarization) task, provides open weights, and can accurately identify the speaker at any moment during a conversation. It supports real-time and recorded audio processing for up to 8 people speaking simultaneously.

Technological breakthroughs and performance

In the stringent VoiceArena speech diarization benchmark test (v1 version), Nemotron-3-Diarization topped the leaderboard with an error rate (DER) of 14.72%, significantly outperforming the system in second place (19.3%). Compared with its predecessor, Streaming Sortformer, the new model reduces the average error rate by 41% across eight test scenarios using an 1.04-second audio buffer. To balance system latency and accuracy, the model supports four levels of dynamic audio buffer settings, ranging from 0.32 seconds to 30.4 seconds.

Functional application layer

The new model can seamlessly work with speech recognition systems such as Parakeet to automatically generate transcripts with anonymized speaker labels (e.g., "speaker_2"). Although the error rate will increase as the number of participants grows, background noise becomes excessive, or reverberation is severe, its powerful underlying architecture can effectively address traditional technical challenges such as overlapping speech.

With this release, NVIDIA introduces a lightweight yet free, high-precision speech diarization model that significantly lowers the technical barrier for complex speech analysis. This move not only provides highly cost-effective underlying support for applications such as real-time meeting transcription, intelligent customer service, and multi-speaker voice interactions, but also indicates that multi-speaker identification technology will be further accelerated in real-world deployment on edge devices and in real-time commercial scenarios.