multivent-raw
143,288 short-video chunks paired with keyframes, OCR text, ASR transcripts,
keyframe captions, embeddings, and full retrieval/claim annotations.
Designed for retrieval and analysis research. Every artifact is per-chunk
and joinable by chunk_id.
At a glance
| Chunks | 143,288 |
| Source videos | 118,791 |
| Total duration | 5,353 h |
| Shards | 667 |
| On-disk size | ≈ 5 TB |
| Topics | 130 |
| Personas | 129 |
| Queries | 222 |
Directory layout
multivent-raw/
├── README.md
│
├── annotations/ ← eval inputs (see Annotations below)
│ ├── personas.jsonl
│ ├── queries.jsonl
│ ├── topics.jsonl
│ ├── reference-claims-{video,event,persona,query}.json
│ ├── {video,event,persona,query}.qrels
│ └── judgments-{video,event,persona,query}.jsonl
│
├── videos/ ← .mp4 + per-chunk JSON
│ ├── catalog.csv
│ └── shard_NNNNNN.tar (×667)
│
├── keyframes/uniform_5s/ ← .jpg frames, one every 5 s
│ ├── catalog.csv
│ └── shard_NNNNNN.tar (×667)
│
├── keyframe-captions/uniform_5s_qwen-multiframe-ic/
│ └── max_imgs_16/ ← per-chunk keyframe captions (.txt; no catalog)
│ └── shard_NNNNNN.tar (×667)
│
├── ocr/
│ ├── ppocrvl15/ ← per-frame OCR text (PaddleOCR-VL-1.5)
│ ├── ppocrv5icdar/ ← per-frame OCR text (PP-OCRv5 ICDAR)
│ └── pagectc/ ← per-frame OCR text (OCR-VL-501 pageCTC)
│ ├── catalog.csv
│ └── shard_NNNNNN.tar (×667)
│
├── asr/
│ ├── qwen3asr1p7b/ ← per-chunk ASR (Qwen3-ASR-1.7B)
│ └── whisperxlargev3/ ← per-chunk ASR (WhisperX/large-v3)
│ ├── catalog.csv
│ └── shard_NNNNNN.tar (×667)
│
└── embeddings/
├── kf_uni5s-vizemb_qwen3vlemb2b/ ← vision embedding, dim 2048
├── kf_uni5s-vizemb_qwen3vlemb8b/ ← vision embedding, dim 4096
├── kf_uni5s-ocr_ppocrvl15-txtemb_qwen3emb8b/ ← text embedding of ppocrvl15 OCR, dim 4096
├── kf_uni5s-ocr_ppocrv5icdar-txtemb_qwen3emb8b/ ← text embedding of ppocrv5icdar OCR, dim 4096
└── kf_uni5s-ocr_pagectc-txtemb_qwen3emb8b/ ← text embedding of pagectc OCR, dim 4096
├── catalog.csv
└── shard_NNNNNN.tar (×667)
Each sharded artifact directory contains one catalog.csv plus the
shard_NNNNNN.tar WebDataset shards (keyframe-captions/ ships shards
only). annotations/ holds plain JSON/JSONL files, not shards.
Identifiers
Three IDs let you locate, group, and time-align everything.
| field | example | what it identifies |
|---|---|---|
chunk_id |
XM5xOIzL_vSkGAKR_0000 |
one chunk; the join key across artifacts |
video_id |
XM5xOIzL_vSkGAKR |
the source video the chunk came from |
frame tNNNNNN |
t000005 |
a keyframe within a chunk, at second NNNNNN of the chunk |
chunk_id == f"{video_id}_{chunk_index:04d}"— always 4-digit padded, even for single-chunk videos.tNNNNNNis the integer second offset within the chunk (zero-padded to 6 digits). Keyframes are sampled every 5 s, so the values aret000000, t000005, t000010, ….- No
chunk_idorvideo_idstarts with-, so filenames are safe to pass totar,find,xargs, etc. without escaping.
Annotations (annotations/)
Retrieval and claim annotations over 130 topics, in the same file set and
schemas as microvent — see the Annotations section of the microvent
dataset's README for field-level documentation of every file. Summary:
annotations/
├── personas.jsonl 129 rows, one per persona
├── queries.jsonl 222 rows, one per query (embeds its persona + topic)
├── topics.jsonl 130 rows, one per topic
├── reference-claims-video.json 5,853 claims — raw, video-centric observations
├── reference-claims-event.json 5,321 claims — event-centric facts
├── reference-claims-persona.json 5,106 claims — the facts each persona cares about
├── reference-claims-query.json 8,310 claims — the facts that answer each query
├── video.qrels 1,365 rows ┐ positive-only qrels,
├── event.qrels 1,365 rows │ one file per claim stage
├── persona.qrels 1,357 rows │ (TREC four-column format)
├── query.qrels 1,303 rows ┘
└── judgments-{video,event,persona,query}.jsonl the same qrels as JSON Lines (adds `language`)
Where microvent varies the persona (some topics get two), multivent-raw
varies the query: each topic has exactly one persona (persona_ids are
UUIDs, and one persona serves two topics, so 130 topics map to 129 distinct
personas), while 92 of the 130 topics carry two queries — one biased and
one unbiased — and the remaining 38 carry only a biased query, for 222
queries in all.
Structurally that means the video, event, and persona claim files have one
claim_set per topic (130), while the query file has one per query (222), the
92 two-query topics contributing two claim_sets apiece. 184 of the query
file's 8,310 claims are negative assertions (meta.is_negative_assertion,
empty evidence, introduced at the query stage).
In-shard file names
Inside every shard, members follow:
<chunk_id>.<artifact_tag>.<extension>
<artifact_tag> matches the artifact directory name (with - → .):
| artifact directory | tag |
|---|---|
videos/ |
(none — videos are the canonical source) |
keyframes/uniform_5s/ |
kf_uni5s |
keyframe-captions/…/max_imgs_16/ |
kf_uni5s.tSSSSSS_tEEEEEE (covered frame range) |
ocr/ppocrvl15/ |
kf_uni5s.ocr_ppocrvl15 |
ocr/ppocrv5icdar/ |
kf_uni5s.ocr_ppocrv5icdar |
ocr/pagectc/ |
kf_uni5s.ocr_pagectc |
asr/qwen3asr1p7b/ |
asr_qwen3asr1p7b |
asr/whisperxlargev3/ |
asr_whisperxlargev3 |
embeddings/kf_uni5s-vizemb_qwen3vlemb2b/ |
kf_uni5s.vizemb_qwen3vlemb2b |
embeddings/kf_uni5s-vizemb_qwen3vlemb8b/ |
kf_uni5s.vizemb_qwen3vlemb8b |
embeddings/kf_uni5s-ocr_ppocrvl15-txtemb_qwen3emb8b/ |
kf_uni5s.ocr_ppocrvl15.txtemb_qwen3emb8b |
embeddings/kf_uni5s-ocr_ppocrv5icdar-txtemb_qwen3emb8b/ |
kf_uni5s.ocr_ppocrv5icdar.txtemb_qwen3emb8b |
embeddings/kf_uni5s-ocr_pagectc-txtemb_qwen3emb8b/ |
kf_uni5s.ocr_pagectc.txtemb_qwen3emb8b |
So:
videos/shard_000000.tar
XM5xOIzL_vSkGAKR_0000.mp4
XM5xOIzL_vSkGAKR_0000.json
keyframes/uniform_5s/shard_000000.tar
XM5xOIzL_vSkGAKR_0000.kf_uni5s.json
XM5xOIzL_vSkGAKR_0000.kf_uni5s.t000000.jpg
XM5xOIzL_vSkGAKR_0000.kf_uni5s.t000005.jpg
…
ocr/ppocrvl15/shard_000000.tar
XM5xOIzL_vSkGAKR_0000.kf_uni5s.ocr_ppocrvl15.jsonl
embeddings/kf_uni5s-vizemb_qwen3vlemb8b/shard_000000.tar
XM5xOIzL_vSkGAKR_0000.kf_uni5s.vizemb_qwen3vlemb8b.npz
embeddings/kf_uni5s-ocr_ppocrvl15-txtemb_qwen3emb8b/shard_000000.tar
XM5xOIzL_vSkGAKR_0000.kf_uni5s.ocr_ppocrvl15.txtemb_qwen3emb8b.npz
If you unpack shard_000000.tar from every artifact into one directory,
the files for a given chunk sort together and don't collide.
The stem before the first . is always the chunk_id — this is what
WebDataset uses to group multi-artifact records for the same chunk into
one sample.
Tracing IDs across artifacts
Pick any artifact and you can walk to any other.
From an OCR record → its keyframe image
The OCR jsonl line:
{"frame": "t000005", "raw": "...", "cleaned": "...", "txt": "..."}
…sits inside ocr/ppocrvl15/shard_NNN.tar under the member
<chunk_id>.kf_uni5s.ocr_ppocrvl15.jsonl. The corresponding image lives
at the same (shard, chunk_id) in the keyframes artifact:
keyframes/uniform_5s/shard_NNN.tar
└── <chunk_id>.kf_uni5s.t000005.jpg ← same `frame` value
The frame field in the OCR line and the tNNNNNN segment of the jpg
filename are the same string. Trivially joinable by membership.
From a keyframe / OCR record → its source video
Locate the chunk in videos/:
videos/shard_NNN.tar
└── <chunk_id>.mp4
└── <chunk_id>.json
shard_NNN is the same number you read the keyframe / OCR from (every
artifact shards identically). The chunk's .json gives you the offset
back into the original full-length video:
{
"video_id": "XM5xOIzL_vSkGAKR",
"chunk_index": 0,
"chunk_count": 1,
"chunk_start_sec": 0.0,
"chunk_end_sec": 36.801,
…
}
The original-video timestamp of frame tNNNNNN is
chunk_start_sec + int(NNNNNN). For single-chunk videos
chunk_start_sec == 0, so the keyframe second is the same in chunk
time and source time.
From an embedding row → its keyframe and OCR record
Inside an embedding npz:
data = np.load("...npz")
data["keyframe_ids"] # array(['t000000', 't000005', ...])
data["embeddings"] # (N, D) float32, L2-normalised
Row i of embeddings describes frame keyframe_ids[i] of chunk_id
(the .npz filename's stem). That same tNNNNNN value indexes the
keyframe jpg and the OCR jsonl line for that frame.
For the text-embedding artifact: keyframe_ids may be a strict subset
of the OCR frame values (frames with empty txt are dropped — see
per-artifact details below).
Finding which shard holds a chunk_id
Every catalog has a shard_index column. One line:
shard = pd.read_csv("multivent-raw/videos/catalog.csv") \
.query(f"chunk_id == '{CID}'")["shard_index"].iat[0]
The answer is the same regardless of which artifact's catalog you check.
Per-artifact details
videos/
Per chunk: one .mp4 (H.264 video, AAC audio where present) and one
.json with the chunk's location in the source video.
Chunk JSON schema:
{
"chunk_id": "XM5xOIzL_vSkGAKR_0000",
"video_id": "XM5xOIzL_vSkGAKR",
"chunk_index": 0,
"duration_sec": 36.801,
"source_duration_sec":36.801,
"chunk_count": 1,
"chunk_start_sec": 0.0,
"chunk_end_sec": 36.801,
"width": 360,
"height": 640,
"fps": 25.0
}
videos/catalog.csv: chunk_id, video_id, chunk_index, chunk_count, shard_index, duration_sec, chunk_start_sec, chunk_end_sec, size_bytes, vcodec, acodec.
keyframes/uniform_5s/
Per chunk: one .kf_uni5s.json (chunk-level metadata, extends the videos
JSON with frame info) and N .kf_uni5s.tNNNNNN.jpg images, one every 5 s.
The chunk JSON adds:
"frame_count": 8,
"frame_period_sec": 5.0,
"frame_timestamps_sec": [0.0, 5.0, 10.0, 15.0, 20.0, 25.0, 30.0, 35.0]
keyframes/uniform_5s/catalog.csv: chunk_id, video_id, chunk_index, shard_index, chunk_count, frame_count, duration_sec.
ocr/ppocrvl15/
Per chunk: one .kf_uni5s.ocr_ppocrvl15.jsonl, with one JSON object per
keyframe (in tNNNNNN order, length == frame_count).
Per-line schema:
{
"frame": "t000000",
"raw": "...model output with <|LOC_NNN|> coordinate tokens...",
"cleaned": "...same content, with repetition-loop artifacts trimmed; LOC tokens preserved...",
"txt": "...LOC tokens stripped, whitespace tidied — ready for grep / text embedders..."
}
Most consumers want txt. Use cleaned if you need the OCR's spatial
layout tokens (each <|LOC_N|> is a coordinate index 0–999). raw is
preserved verbatim for anyone who wants pre-cleanup output.
ocr/ppocrvl15/catalog.csv: chunk_id, video_id, chunk_index, shard_index, n_frames, n_frames_repetition_cleaned.
asr/{qwen3asr1p7b,whisperxlargev3}/
Per chunk: one .asr_<backend>.json with language detection, voice-activity
detection, speaker diarization, and a transcript. Both backends share one
wrapper schema (abridged):
{
"chunk_id": "XM5xOIzL_vSkGAKR_0000",
"video_id": "XM5xOIzL_vSkGAKR",
"chunk_index": 0,
"duration_s": 38.824,
"language": {"hint": null, "detected": "Russian", "from_dir": false},
"asr_backend": "qwen", // or "whisperx"
"vad": [{"start": 1.19, "end": 4.05}, ...],
"diarization": [{"start": 1.19, "end": 4.05, "speaker_id": "SPEAKER_00"}, ...],
"overlap": [...],
"embeddings": [{"speaker_id": "SPEAKER_00", "vector": [...], "n_turns": 3}, ...],
"transcript": {
"timestamp_resolution": "segment", // "word" for whisperx
"segments": [{"start": 1.19, "end": 4.05, "text": "...", "speaker_id": "SPEAKER_00", "flags": []}, ...],
"words": [] // whisperx fills this with word-level timings
},
"qc_flags": []
}
The chunk's transcript text is the concatenation of
transcript["segments"][*]["text"]. qwen3asr1p7b produces segment-level
timestamps; whisperxlargev3 additionally fills transcript["words"] with
word-level timings. The embeddings field holds per-speaker voice
embeddings from diarization — unrelated to the embeddings/ artifact
directory.
keyframe-captions/uniform_5s_qwen-multiframe-ic/max_imgs_16/
Per chunk: one or more .txt members, each an English prose caption of a
window of up to 16 consecutive keyframes, produced by Qwen multi-frame image
captioning. The member name encodes the covered frame range:
<chunk_id>.kf_uni5s.t000000_t000075.txt ← caption of frames t000000 … t000075
Chunks with ≤ 16 keyframes get a single member covering the whole chunk;
longer chunks get one member per 16-keyframe window. No catalog.csv for
this artifact — enumerate members from the tars, or take chunk lists from
videos/catalog.csv.
embeddings/kf_uni5s-vizemb_qwen3vlemb{2b,8b}/ (vision)
Per chunk: one .npz with the keyframes encoded by Qwen3-VL-Embedding
(2B or 8B).
| key | shape | dtype |
|---|---|---|
keyframe_ids |
(N,) |
<U7 |
embeddings |
(N, D) |
float32 |
D = 2048 for the 2B model, D = 4096 for the 8B model. Row i
describes frame keyframe_ids[i]. Embeddings are L2-normalised, so
cosine similarity == inner product.
catalog.csv: chunk_id, video_id, chunk_index, shard_index, n_frames, dim.
embeddings/kf_uni5s-ocr_{ppocrvl15,ppocrv5icdar,pagectc}-txtemb_qwen3emb8b/ (text)
Per chunk: one .npz with the OCR text of each keyframe encoded by
Qwen3-Embedding-8B. Three sibling artifacts, one per OCR backend.
| key | shape | dtype |
|---|---|---|
keyframe_ids |
(M,) |
<U7 |
embeddings |
(M, 4096) |
float32 |
M ≤ N. Frames whose source-OCR text was empty are skipped (no text →
nothing to embed). Chunks with no text in any frame produce no member in
the tar at all; their catalog row has n_frames_embedded == 0. The
zero-coverage rate varies by OCR backend:
| backend | source field | zero-emb chunks |
|---|---|---|
ppocrvl15 |
txt |
≈ 11 % |
ppocrv5icdar |
text |
≈ 8 % |
pagectc |
text |
≈ 3 % |
catalog.csv (same shape for all three): chunk_id, video_id, chunk_index, shard_index, n_frames_embedded, n_frames_skipped, dim.
Catalogs
One catalog.csv per sharded artifact directory (keyframe-captions/ is
the exception — it has none), 143,288 rows each. They all share the
canonical prefix (chunk_id, video_id, chunk_index, shard_index), so any
pair joins on chunk_id:
import pandas as pd
videos = pd.read_csv("multivent-raw/videos/catalog.csv")
ocr = pd.read_csv("multivent-raw/ocr/ppocrvl15/catalog.csv")
txtemb = pd.read_csv("multivent-raw/embeddings/kf_uni5s-ocr_ppocrvl15-txtemb_qwen3emb8b/catalog.csv")
df = videos.merge(ocr, on="chunk_id").merge(txtemb, on="chunk_id",
suffixes=("_ocr", "_txtemb"))
# guaranteed 143,288 rows
Loading examples
Plain tarfile
import tarfile, json
with tarfile.open("multivent-raw/ocr/ppocrvl15/shard_000000.tar") as tf:
for m in tf:
if not m.name.endswith(".jsonl"):
continue
for line in tf.extractfile(m).read().decode().splitlines():
rec = json.loads(line)
print(rec["frame"], rec["txt"][:60])
break
One chunk's npz, by name
import io, tarfile, numpy as np
CID = "XM5xOIzL_vSkGAKR_0000"
SHARD = 0
TAR = f"multivent-raw/embeddings/kf_uni5s-vizemb_qwen3vlemb8b/shard_{SHARD:06d}.tar"
MEMBER = f"{CID}.kf_uni5s.vizemb_qwen3vlemb8b.npz"
with tarfile.open(TAR) as tf:
blob = tf.extractfile(MEMBER).read()
data = np.load(io.BytesIO(blob))
print(data["keyframe_ids"]) # ['t000000' 't000005' ...]
print(data["embeddings"].shape) # (N, 4096)
WebDataset (multi-artifact, joined by chunk_id)
import webdataset as wds
url = "multivent-raw/{videos,embeddings/kf_uni5s-ocr_ppocrvl15-txtemb_qwen3emb8b}/shard_000000.tar"
ds = wds.WebDataset(url, shardshuffle=False).decode()
for sample in ds:
chunk_id = sample["__key__"]
mp4_bytes = sample.get("mp4")
txt_emb = sample.get("kf_uni5s.ocr_ppocrvl15.txtemb_qwen3emb8b.npz")
...
Sharding
667 shards of ~210 chunks each (range 121–272). A chunk lives in exactly
one shard across all artifacts: shard 42 of videos/ and shard 42 of any
other artifact describe the same set of chunks. Shards are independent,
so you can shuffle shard order then iterate within for IID-ish batches.
wds.WebDataset(shardshuffle=True) does this automatically.
- Downloads last month
- 34