Hugging Face released the Vaani Noise Event Timestamp Dataset on September 7, 2026, saying it provides precise temporal localization of noise events in real-world speech. It is the company's first noise-robust speech dataset since its earlier multilingual speech project, Project Vaani.

Hugging Face reported 122+ hours of total audio, with 72,756 speech segments and 38,541 speakers, covering 58 Indian languages. That compares with its prior dataset, which focused on clean speech without noise annotations.

The Vaani Noise Event Timestamp Dataset is built on Project Vaani's spontaneous multilingual speech and targets noise-robust speech recognition. Availability begins with the full dataset on Hugging Face, initially for researchers and developers.

"Speech AI has become remarkably good at hearing us, as long as we speak in a quiet room," said Suryansh Shukla, a contributor to the project. The dataset is designed to address the limitation that accuracy drops when speech AI meets real-life noise.

The announcement follows the release of Project Vaani, which provided clean speech data for multiple Indian languages. Hugging Face said the new dataset fills a gap by combining real, co-occurring background noise with natural speech.

Hugging Face did not say how the dataset will be used beyond research, and raised the open question of how to scale noise-robustness research to low-resource languages. The dataset will be used to stress-test voice AI systems in real-world Indian conditions.

Source: huggingface