Hugging Face released Open Yap 1K on September 3, 2026, saying it provides 1,000 hours of full-duplex English conversation for free commercial use. It is the company's first major dataset update since its earlier speech corpora releases.
Hugging Face reported 1,000 hours of recorded data, measured at 48kHz, with 1,602 conversations captured in real-world environments. That compares with smaller datasets like Fisher English, which contains 1,959 hours of telephone-band speech.
Open Yap 1K is built on a WhatsApp-like app that captures natural conversations between friends and family, targeting full-duplex speech models. Availability begins with a sample of 8.9 hours on the Hugging Face Hub, initially for researchers and developers.
"We built a WhatsApp-like app that people used instead of their regular phone calls to talk to friends and family," said Christian Vestergaard, a contributor. The dataset aims to model real conversational dynamics, including interruptions and laughter.
The announcement follows Hugging Face's earlier focus on speech corpora like Fisher and Switchboard. The company framed the significance of Open Yap 1K as a step toward more natural and realistic speech modeling.
Hugging Face did not say how the dataset will be used in commercial applications, and raised questions about the challenges of modeling real-time speech interactions. The company said the dataset is available on request under the Open Yap 1K Data Use Agreement.
Source: huggingface