The dataset viewer is not available for this split.
Error code: TooBigContentError
Need help to make the dataset viewer work? Make sure to review how to configure the dataset viewer, and open a discussion for direct support.
YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
YouTube Commons and EUVoxCommons
Datasets Documentation
Executive Summary
This datacard provides a comprehensive description of YouTube-Commons and EUVoxCommons (European Parliament proceedings) collected and handled by pleias along with a sample of 1,304 documented audio files.
These datasets represent the largest collection of fully open-source copyright-compliant speech data for the 24 official languages of the European Union and more.
Key Statistics:
- YouTube-Commons: 1,094,136 hours (labeled and unlabeled)
- EUVoxCommons: 387,082 hours (385,291 unlabeled + 1,791 labeled)
- Languages covered: 23 out of 24 EU languages (all except Irish with minimal data) + Asian and African languages
- Licensing: CC-BY 3.0 (YouTube-Commons), CC-0 (EUVoxCommons)
- Pseudo-labels provided: 441,206 hours of automatic transcriptions
The combined YouTube-Commons and VoxPopuli dataset covers 23 out of 24 official EU languages, as well as multiple other Asian languages.
Sample description
This datacard features a sample of 1,136 audio files from YouTube (394 hours) published from 2009 to 2026 in 93 languages and 168 audio files from VoxPopuli/Europarl (84 hours) published from 2009–2020.
Samples use the same structure as the final dataset and are distributed as 12 fully shuffled parquet files with three components:
- Metadata scraped from the Youtube official API with original url, channel information, date of publication.
- New transcripts made with state of the art ASR model, not the Youtube transcript (frequently faulty). These allow for full text search of the entire corpus.
- Audio included directly in the parquet file, as is now common expectation for multimodal model training.
Samples are representative of the entire corpus, especially in terms of multilingual diversity (roughly half English, then Spanish, French, Italian, Korean, German…)
We intently selected for longer samples, as they make up a larger share of the total runtime in the full corpus.
Comparison with other datasets
| Dataset/Model | Hours | Languages | OS-Compliant | Training Data Public |
|---|---|---|---|---|
| YouTube-Commons | 1,094,136 hours | 18 EU + 180+ other languages | ✓ Yes | ✓ Yes |
| VoxPopuli | 387,082 | 23 EU | ✓ Yes | ✓ Yes |
| Whisper v2 training | ~680,000 | 99 | ✗ No | ✗ No |
| Whisper v3 training | ~5,000,000 | 99 | ✗ No | ✗ No |
| OWSM training | ~180,000 | 151 | ✗ No | ✓ Yes (but not OS-compliant) |
| MLS (open-source) | 50,687 | 8 | ✓ Yes | ✓ Yes |
| Common Voice | 6,732 | 22 EU | ✓ Yes | ✓ Yes |
| GigaSpeech | 33,000 | 1 | ✗ No | ✓ Yes (but not OS-compliant) |
Technical Specifications
Audio formats:
- VoxPopuli: Primarily MP3, some WAV
- YouTube-Commons: Variable (MP3, M4A, WebM audio)
- Sampling rates: Typically 16kHz
- Bit depth: 16-bit or variable
- Channels: Mono or stereo (convertible)
Transcript formats:
- Plain text (UTF-8 encoding)
- JSON/JSONL with metadata
- Alignment information where available
YouTube-Commons Dataset
Overview
YouTube-Commons is a large-scale multilingual speech corpus extracted from YouTube videos released under CC-BY licenses.
Structure
Youtube-Commons is currently composed of three different subsets collected at different times. The first subset is immediately available, Youtube-Commons-2 is under processing and Youtube-Commons-3 getting collected. All subsets are multilingual (40% English).
| Subset | Volume |
|---|---|
| Youtube-Commons-1 | 287,273 hours (1,391,708 audio samples) |
| Youtube-Commons-2 | 159,384 hours (772143 audio samples) |
| Youtube-Commons-3* | up to 647,479 hours (2,199,311 audio samples) |
*Estimate
In total, the final Youtube-Commons dataset is 4,363,162 videos (1,094,136 hours).
Language Coverage
| Language | Hours | Share |
|---|---|---|
| English | 734,177 | 67.1% |
| Spanish | 66,895 | 6.1% |
| French | 54,521 | 5.0% |
| Russian | 42,265 | 3.9% |
| Portuguese | 29,350 | 2.7% |
| German | 25,940 | 2.4% |
| Korean | 19,874 | 1.8% |
| Italian | 18,077 | 1.7% |
| Indonesian | 15,335 | 1.4% |
| Hindi | 13,156 | 1.2% |
| Vietnamese | 11,275 | 1.0% |
| Japanese | 6,085 | 0.6% |
| Turkish | 4,956 | 0.5% |
| Dutch | 4,544 | 0.4% |
| Urdu | 2,921 | 0.3% |
| Arabic | 2,878 | 0.3% |
| Bengali | 2,591 | 0.2% |
| Tamil | 2,239 | 0.2% |
| Polish | 2,056 | 0.2% |
| Chinese | 1,916 | 0.2% |
| Ukrainian | 1,447 | 0.1% |
| Telugu | 1,347 | 0.1% |
| Punjabi | 1,337 | 0.1% |
| Thai | 1,329 | 0.1% |
| Malayalam | 1,173 | 0.1% |
| Catalan | 979 | 0.1% |
| Kannada | 943 | 0.1% |
| Malay | 781 | 0.1% |
| Filipino | 743 | 0.1% |
| Other languages (166) | 23,006 | 2.1% |
| Total | 1,094,136 | 100% |
Data Characteristics
Content domains:
- Educational content
- Entertainment and media
- News and documentaries
- Vlogs and personal content
- Lectures and presentations
- Various other YouTube content categories
Audio characteristics:
- Variable recording quality (user-generated content)
- Multiple speakers per video
- Background music and noise present in many samples
- Natural, conversational speech styles
- Mixed acoustic environments
EUVoxCommons
Overview
VoxPopuli is derived from European Parliament event recordings, providing parliamentary speech data across EU languages.
Dataset specifications:
- Total volume: 387,082 hours
- Unlabeled data: 385,291 hours
- Labeled data: 1,791 hours
- License: CC-0 (Public Domain)
- Source: European Parliament recordings
- Domain: Parliamentary proceedings
Comprehensive Language Coverage
It provides substantial data for 23 EU languages:
| Language | Unlabeled (hours) | Labeled (hours) | Total (hours) |
|---|---|---|---|
| Bulgarian (bg) | 17,609 | - | 17,609 |
| Croatian (hr) | 8,106 | 55 | 8,161 |
| Czech (cs) | 18,705 | - | 18,705 |
| Danish (da) | 13,600 | - | 13,600 |
| Dutch (nl) | 19,014 | - | 19,014 |
| English (en) | 84,704 | - | 84,704 |
| Estonian (et) | 10,604 | - | 10,604 |
| Finnish (fi) | 14,200 | - | 14,200 |
| French (fr) | 22,896 | - | 22,896 |
| German (de) | 23,228 | - | 23,228 |
| Greek (el) | 17,703 | - | 17,703 |
| Hungarian (hu) | 17,701 | - | 17,701 |
| Italian (it) | 21,933 | - | 21,933 |
| Latvian (lv) | 13,100 | - | 13,100 |
| Lithuanian (lt) | 14,400 | - | 14,400 |
| Maltese (mt) | 9,100 | - | 9,100 |
| Polish (pl) | 21,207 | - | 21,207 |
| Portuguese (pt) | 17,526 | - | 17,526 |
| Romanian (ro) | 17,906 | - | 17,906 |
| Slovak (sk) | 12,100 | - | 12,100 |
| Slovenian (sl) | 11,300 | - | 11,300 |
| Spanish (es) | 21,526 | - | 21,526 |
| Swedish (sv) | 16,300 | - | 16,300 |
| Total | 385,291 | 1,791 | 387,082 |
Key characteristics:
- Minimum 8,000+ hours for all languages except Irish (no data available)
- More balanced distribution compared to YouTube-Commons
- Consistent data volume across medium and large languages
Data Characteristics
Recording quality:
- Professional studio recordings
- High-quality microphones and equipment
- Controlled acoustic environments
- Minimal background noise
Speech characteristics:
- Formal register (parliamentary proceedings)
- Prepared and spontaneous speech
- Multiple speakers per session
- Native and non-native speakers
- Various accents and dialects
Content domains:
- Political discourse
- Legislative discussions
- Policy debates
- Official statements and speeches
- Committee proceedings
Linguistic features:
- Formal vocabulary
- Technical and legal terminology
- Complex sentence structures
- Code-switching (multilingual speakers)
- Proper names and institutional references
- Downloads last month
- -