NVIDIA released Nemotron 3 Diarization on September 23, 2026, saying it enables real-time, multi-speaker identification in conversations. It is the company's first diarization update since the introduction of the Streaming Sortformer model.
NVIDIA reported a 14.72% Diarization Error Rate (DER), measured on VoiceArena's Diarization-Bench leaderboard. That compares with a 11.19% DER from earlier models.
Nemotron 3 Diarization is built on a 100M-parameter model and targets applications requiring speaker attribution in overlapping conversations. Availability begins with open-weight access, initially for developers and researchers.
"Turn overlapping conversations into speaker-aware data with one open-weight, 100M-parameter model," said Francesco fciannella. The model handles up to eight speakers and supports flexible recording lengths.
The announcement follows NVIDIA's earlier work on Streaming Sortformer. NVIDIA said the model expands support to eight speakers and improves accuracy and throughput.
NVIDIA did not say how the model performs in noisy environments, and it raises the question of how it handles domain shifts. The model supports as many as eight speaker channels.
Source: huggingface