
Speech models need more than clean studio recordings if they are to work in noisy rooms, with different accents, and across many languages. The ESPnet research team released YODAS v3 on September 27: an openly available dataset containing more than 1.1 million hours of audio. The new collection could help researchers develop speech technology outside the large proprietary data holdings of major companies.
Key takeaways
- YODAS v3 contains more than 1.1 million hours of web audio and is distributed under CC BY 3.0.
- The files use a 48 kHz format; more than 70 percent contain at least two genuinely distinct audio channels.
- Transcripts and English translations cover only part of the collection and may contain errors.
- The dataset supports research into speech recognition, translation, and audio processing, but it is not a finished voice assistant.
What makes this collection different
The team collected videos released under a Creative Commons license and preserved their audio in its original Opus format. Earlier large speech collections were often reduced to one channel and lower sample rates. That can be sufficient for recognizing spoken words. Researchers studying spatial sound, multiple speakers, or speech synthesis often need more of the original recording. YODAS v3 is therefore more than a larger text collection with audio attached.
The researchers report that more than 70 percent of recordings have two distinguishable channels. That distinction matters: a file with two channels is not necessarily true stereo, because both channels could be identical. Likewise, a 48 kHz label does not guarantee equally detailed sound. According to the dataset card, more than 92 percent of recordings reach an estimated effective sample rate of at least 32 kHz, while about 67 percent reach 44.1 kHz. Researchers can use that metadata to select material suited to their task.
Many languages, imperfect labels
The paper lists 147 languages, but that figure partly relies on the language or locale supplied by the video platform. The more detailed dataset card now organizes files according to the language of available transcripts and lists 102 language codes plus a folder for recordings without transcripts. Those counts reflect different assignment methods. Anyone assessing coverage for a training project should inspect the metadata and usable hours for each language instead of relying on the headline number alone.
About three quarters of the data has transcripts. English translations exist for roughly 60 percent of the non-English material. This creates opportunities for subtitling and speech translation research, but it does not mean every text was checked by a human. The paper explicitly describes the collection as weakly labeled: many captions come from automatic recognition. Language IDs apply to whole files and may be wrong for individual sentences. Smaller languages could benefit from the scale without gaining perfectly reliable labels. That tension also appears in our look at AI for smaller languages.
The selection of recordings also deserves attention. The team searched for videos using language-specific terms and favored newer uploads. That approach can find more material beyond the largest languages, but it cannot represent every speaker or speaking situation equally. The license listed on the dataset page also does not replace checking whether a particular subset suits a planned project. For a model meant to understand phone calls, similarity between its training recordings and actual calls matters more than the total number of hours alone. A large dataset is a starting point for selection and measurement, not a substitute for either.
What the first experiments show
The authors do not claim to have built a production voice assistant from the entire collection. Among other experiments, they trained speech recognition models on selected language subsets and compared different filters for the transcripts. In this limited setup, less aggressive filtering improved error rates for most languages studied. That suggests the additional data was useful under those conditions. It does not establish equal quality across all languages or superiority over commercial systems.
A second experiment examined audio codecs, which turn sound into compact representations for storage and transmission. A codec trained on high-resolution YODAS data performed better in the reported tests than the lower sample rate variants used in that comparison. Those results cover particular datasets and metrics; they do not prove audible quality in every real application. Keeping the research result separate from a product promise is essential to understanding the release.
Outlook: openness still requires careful selection
The dataset is publicly available on Hugging Face, which lists its size as 55.4 terabytes. That is not a casual download for an individual. Researchers can examine subsets by language, transcript availability, effective audio bandwidth, and channel count. Whether the collection leads to better speech tools for underserved languages will then depend on verified training data, meaningful evaluations, and actual deployment conditions. YODAS v3 provides raw material at a rarely seen scale, not a ready-made verdict on quality.

