Dataset Viewer
Auto-converted to Parquet Duplicate
The dataset viewer is not available for this split.
Parquet error: Scan size limit exceeded: attempted to read 637040845 bytes, limit is 300000000 bytes Make sure that 1. the Parquet files contain a page index to enable random access without loading entire row groups2. otherwise use smaller row-group sizes when serializing the Parquet files
Error code:   TooBigContentError

Need help to make the dataset viewer work? Make sure to review how to configure the dataset viewer, and open a discussion for direct support.

YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

YouTube Commons and EUVoxCommons

Datasets Documentation

Executive Summary

This datacard provides a comprehensive description of YouTube-Commons and EUVoxCommons (European Parliament proceedings) collected and handled by pleias along with a sample of 1,304 documented audio files.

These datasets represent the largest collection of fully open-source copyright-compliant speech data for the 24 official languages of the European Union and more.

Key Statistics:

  • YouTube-Commons: 1,094,136 hours (labeled and unlabeled)
  • EUVoxCommons: 387,082 hours (385,291 unlabeled + 1,791 labeled)
  • Languages covered: 23 out of 24 EU languages (all except Irish with minimal data) + Asian and African languages
  • Licensing: CC-BY 3.0 (YouTube-Commons), CC-0 (EUVoxCommons)
  • Pseudo-labels provided: 441,206 hours of automatic transcriptions

The combined YouTube-Commons and VoxPopuli dataset covers 23 out of 24 official EU languages, as well as multiple other Asian languages.

Sample description

This datacard features a sample of 1,136 audio files from YouTube (394 hours) published from 2009 to 2026 in 93 languages and 168 audio files from VoxPopuli/Europarl (84 hours) published from 2009–2020.

Samples use the same structure as the final dataset and are distributed as 12 fully shuffled parquet files with three components:

  • Metadata scraped from the Youtube official API with original url, channel information, date of publication.
  • New transcripts made with state of the art ASR model, not the Youtube transcript (frequently faulty). These allow for full text search of the entire corpus.
  • Audio included directly in the parquet file, as is now common expectation for multimodal model training.

Samples are representative of the entire corpus, especially in terms of multilingual diversity (roughly half English, then Spanish, French, Italian, Korean, German…)

We intently selected for longer samples, as they make up a larger share of the total runtime in the full corpus.

Comparison with other datasets

Dataset/Model Hours Languages OS-Compliant Training Data Public
YouTube-Commons 1,094,136 hours 18 EU + 180+ other languages ✓ Yes ✓ Yes
VoxPopuli 387,082 23 EU ✓ Yes ✓ Yes
Whisper v2 training ~680,000 99 ✗ No ✗ No
Whisper v3 training ~5,000,000 99 ✗ No ✗ No
OWSM training ~180,000 151 ✗ No ✓ Yes (but not OS-compliant)
MLS (open-source) 50,687 8 ✓ Yes ✓ Yes
Common Voice 6,732 22 EU ✓ Yes ✓ Yes
GigaSpeech 33,000 1 ✗ No ✓ Yes (but not OS-compliant)

Technical Specifications

Audio formats:

  • VoxPopuli: Primarily MP3, some WAV
  • YouTube-Commons: Variable (MP3, M4A, WebM audio)
  • Sampling rates: Typically 16kHz
  • Bit depth: 16-bit or variable
  • Channels: Mono or stereo (convertible)

Transcript formats:

  • Plain text (UTF-8 encoding)
  • JSON/JSONL with metadata
  • Alignment information where available

YouTube-Commons Dataset

Overview

YouTube-Commons is a large-scale multilingual speech corpus extracted from YouTube videos released under CC-BY licenses.

Structure

Youtube-Commons is currently composed of three different subsets collected at different times. The first subset is immediately available, Youtube-Commons-2 is under processing and Youtube-Commons-3 getting collected. All subsets are multilingual (40% English).

Subset Volume
Youtube-Commons-1 287,273 hours (1,391,708 audio samples)
Youtube-Commons-2 159,384 hours (772143 audio samples)
Youtube-Commons-3* up to 647,479 hours (2,199,311 audio samples)

*Estimate

In total, the final Youtube-Commons dataset is 4,363,162 videos (1,094,136 hours).

Language Coverage

Language Hours Share
English 734,177 67.1%
Spanish 66,895 6.1%
French 54,521 5.0%
Russian 42,265 3.9%
Portuguese 29,350 2.7%
German 25,940 2.4%
Korean 19,874 1.8%
Italian 18,077 1.7%
Indonesian 15,335 1.4%
Hindi 13,156 1.2%
Vietnamese 11,275 1.0%
Japanese 6,085 0.6%
Turkish 4,956 0.5%
Dutch 4,544 0.4%
Urdu 2,921 0.3%
Arabic 2,878 0.3%
Bengali 2,591 0.2%
Tamil 2,239 0.2%
Polish 2,056 0.2%
Chinese 1,916 0.2%
Ukrainian 1,447 0.1%
Telugu 1,347 0.1%
Punjabi 1,337 0.1%
Thai 1,329 0.1%
Malayalam 1,173 0.1%
Catalan 979 0.1%
Kannada 943 0.1%
Malay 781 0.1%
Filipino 743 0.1%
Other languages (166) 23,006 2.1%
Total 1,094,136 100%

Data Characteristics

Content domains:

  • Educational content
  • Entertainment and media
  • News and documentaries
  • Vlogs and personal content
  • Lectures and presentations
  • Various other YouTube content categories

Audio characteristics:

  • Variable recording quality (user-generated content)
  • Multiple speakers per video
  • Background music and noise present in many samples
  • Natural, conversational speech styles
  • Mixed acoustic environments

EUVoxCommons

Overview

VoxPopuli is derived from European Parliament event recordings, providing parliamentary speech data across EU languages.

Dataset specifications:

  • Total volume: 387,082 hours
  • Unlabeled data: 385,291 hours
  • Labeled data: 1,791 hours
  • License: CC-0 (Public Domain)
  • Source: European Parliament recordings
  • Domain: Parliamentary proceedings

Comprehensive Language Coverage

It provides substantial data for 23 EU languages:

Language Unlabeled (hours) Labeled (hours) Total (hours)
Bulgarian (bg) 17,609 - 17,609
Croatian (hr) 8,106 55 8,161
Czech (cs) 18,705 - 18,705
Danish (da) 13,600 - 13,600
Dutch (nl) 19,014 - 19,014
English (en) 84,704 - 84,704
Estonian (et) 10,604 - 10,604
Finnish (fi) 14,200 - 14,200
French (fr) 22,896 - 22,896
German (de) 23,228 - 23,228
Greek (el) 17,703 - 17,703
Hungarian (hu) 17,701 - 17,701
Italian (it) 21,933 - 21,933
Latvian (lv) 13,100 - 13,100
Lithuanian (lt) 14,400 - 14,400
Maltese (mt) 9,100 - 9,100
Polish (pl) 21,207 - 21,207
Portuguese (pt) 17,526 - 17,526
Romanian (ro) 17,906 - 17,906
Slovak (sk) 12,100 - 12,100
Slovenian (sl) 11,300 - 11,300
Spanish (es) 21,526 - 21,526
Swedish (sv) 16,300 - 16,300
Total 385,291 1,791 387,082

Key characteristics:

  • Minimum 8,000+ hours for all languages except Irish (no data available)
  • More balanced distribution compared to YouTube-Commons
  • Consistent data volume across medium and large languages

Data Characteristics

Recording quality:

  • Professional studio recordings
  • High-quality microphones and equipment
  • Controlled acoustic environments
  • Minimal background noise

Speech characteristics:

  • Formal register (parliamentary proceedings)
  • Prepared and spontaneous speech
  • Multiple speakers per session
  • Native and non-native speakers
  • Various accents and dialects

Content domains:

  • Political discourse
  • Legislative discussions
  • Policy debates
  • Official statements and speeches
  • Committee proceedings

Linguistic features:

  • Formal vocabulary
  • Technical and legal terminology
  • Complex sentence structures
  • Code-switching (multilingual speakers)
  • Proper names and institutional references
Downloads last month
-