HuggingFace released YODAS v3 on September 27, 2026, saying it triples the scale of its previous dataset while improving quality and usability. It is the company's first major update to the YODAS family since the original release almost 3 years ago.
HuggingFace reported 1.1 million hours of audio, measured on the scale of the YODAS family, compared with ~370,000 hours in YODAS v1. That represents a threefold increase in data volume.
YODAS v3 is built on a multilingual speech dataset and targets research in voice AI, speech recognition, and natural language processing. Availability begins with the full dataset on HuggingFace, initially for the open research community.
"YODAS v3 is the largest open release in the YODAS family to date," said Shinji Watanabe, a researcher at the University of Tokyo. The dataset includes 100+ languages and supports a wide range of speech research tasks.
The announcement follows the release of YODAS v2 and its derived datasets, which have been widely used in the development of production-scale models. HuggingFace did not say when the next version of YODAS will be released, and noted the ongoing challenge of maintaining data diversity across languages.
Source: huggingface