Nvidia released Nemotron 3 Diarization on September 27, 2026, saying it identifies which speaker is talking at any given moment in a conversation. It is the company's first speaker identification update since the release of Streaming Sortformer.
Nvidia reported a DER of 14.7% on the VoiceArena Diarization Benchmark v1, measured on audio recordings with overlapping speech. That compares with a 19.3% error rate from the next best system.
Nemotron 3 Diarization is built on a speech recognition system like Parakeet and targets real-time audio analysis. Availability begins with the model's weights freely available, initially for developers and researchers.
"The model can tell apart up to eight speakers and detect when multiple people talk at the same time," said Jonathan Kemper, the article's author. The model works with both recordings and live audio, though it produces anonymous speaker labels.
The announcement follows the release of Streaming Sortformer. Nvidia framed the significance of the new model as an improvement in accuracy and real-time speaker identification capabilities.
Nvidia did not say how the model will be integrated into existing systems, and it raised the question of whether the model can handle more than eight speakers. The company noted the model's performance will be further tested in upcoming scenarios.
Source: thedecoder