Datasets:

License:

You need to agree to share your contact information to access this dataset

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this dataset content.

multivent-raw

143,288 short-video chunks paired with keyframes, OCR text, ASR transcripts, keyframe captions, embeddings, and full retrieval/claim annotations. Designed for retrieval and analysis research. Every artifact is per-chunk and joinable by chunk_id.


At a glance

Chunks 143,288
Source videos 118,791
Total duration 5,353 h
Shards 667
On-disk size ≈ 5 TB
Topics 130
Personas 129
Queries 222

Directory layout

multivent-raw/
├── README.md
│
├── annotations/                                              ← eval inputs (see Annotations below)
│   ├── personas.jsonl
│   ├── queries.jsonl
│   ├── topics.jsonl
│   ├── reference-claims-{video,event,persona,query}.json
│   ├── {video,event,persona,query}.qrels
│   └── judgments-{video,event,persona,query}.jsonl
│
├── videos/                                                   ← .mp4 + per-chunk JSON
│   ├── catalog.csv
│   └── shard_NNNNNN.tar   (×667)
│
├── keyframes/uniform_5s/                                     ← .jpg frames, one every 5 s
│   ├── catalog.csv
│   └── shard_NNNNNN.tar   (×667)
│
├── keyframe-captions/uniform_5s_qwen-multiframe-ic/
│   └── max_imgs_16/                                          ← per-chunk keyframe captions (.txt; no catalog)
│       └── shard_NNNNNN.tar   (×667)
│
├── ocr/
│   ├── ppocrvl15/                                            ← per-frame OCR text (PaddleOCR-VL-1.5)
│   ├── ppocrv5icdar/                                         ← per-frame OCR text (PP-OCRv5 ICDAR)
│   └── pagectc/                                              ← per-frame OCR text (OCR-VL-501 pageCTC)
│       ├── catalog.csv
│       └── shard_NNNNNN.tar   (×667)
│
├── asr/
│   ├── qwen3asr1p7b/                                         ← per-chunk ASR (Qwen3-ASR-1.7B)
│   └── whisperxlargev3/                                      ← per-chunk ASR (WhisperX/large-v3)
│       ├── catalog.csv
│       └── shard_NNNNNN.tar   (×667)
│
└── embeddings/
    ├── kf_uni5s-vizemb_qwen3vlemb2b/                         ← vision embedding, dim 2048
    ├── kf_uni5s-vizemb_qwen3vlemb8b/                         ← vision embedding, dim 4096
    ├── kf_uni5s-ocr_ppocrvl15-txtemb_qwen3emb8b/             ← text embedding of ppocrvl15 OCR,    dim 4096
    ├── kf_uni5s-ocr_ppocrv5icdar-txtemb_qwen3emb8b/          ← text embedding of ppocrv5icdar OCR, dim 4096
    └── kf_uni5s-ocr_pagectc-txtemb_qwen3emb8b/               ← text embedding of pagectc OCR,      dim 4096
        ├── catalog.csv
        └── shard_NNNNNN.tar   (×667)

Each sharded artifact directory contains one catalog.csv plus the shard_NNNNNN.tar WebDataset shards (keyframe-captions/ ships shards only). annotations/ holds plain JSON/JSONL files, not shards.


Identifiers

Three IDs let you locate, group, and time-align everything.

field example what it identifies
chunk_id XM5xOIzL_vSkGAKR_0000 one chunk; the join key across artifacts
video_id XM5xOIzL_vSkGAKR the source video the chunk came from
frame tNNNNNN t000005 a keyframe within a chunk, at second NNNNNN of the chunk
  • chunk_id == f"{video_id}_{chunk_index:04d}" — always 4-digit padded, even for single-chunk videos.
  • tNNNNNN is the integer second offset within the chunk (zero-padded to 6 digits). Keyframes are sampled every 5 s, so the values are t000000, t000005, t000010, ….
  • No chunk_id or video_id starts with -, so filenames are safe to pass to tar, find, xargs, etc. without escaping.

Annotations (annotations/)

Retrieval and claim annotations over 130 topics, in the same file set and schemas as microvent — see the Annotations section of the microvent dataset's README for field-level documentation of every file. Summary:

annotations/
├── personas.jsonl                        129 rows, one per persona
├── queries.jsonl                         222 rows, one per query (embeds its persona + topic)
├── topics.jsonl                          130 rows, one per topic
├── reference-claims-video.json           5,853 claims — raw, video-centric observations
├── reference-claims-event.json           5,321 claims — event-centric facts
├── reference-claims-persona.json         5,106 claims — the facts each persona cares about
├── reference-claims-query.json           8,310 claims — the facts that answer each query
├── video.qrels                           1,365 rows ┐ positive-only qrels,
├── event.qrels                           1,365 rows │ one file per claim stage
├── persona.qrels                         1,357 rows │ (TREC four-column format)
├── query.qrels                           1,303 rows ┘
└── judgments-{video,event,persona,query}.jsonl     the same qrels as JSON Lines (adds `language`)

Where microvent varies the persona (some topics get two), multivent-raw varies the query: each topic has exactly one persona (persona_ids are UUIDs, and one persona serves two topics, so 130 topics map to 129 distinct personas), while 92 of the 130 topics carry two queries — one biased and one unbiased — and the remaining 38 carry only a biased query, for 222 queries in all.

Structurally that means the video, event, and persona claim files have one claim_set per topic (130), while the query file has one per query (222), the 92 two-query topics contributing two claim_sets apiece. 184 of the query file's 8,310 claims are negative assertions (meta.is_negative_assertion, empty evidence, introduced at the query stage).


In-shard file names

Inside every shard, members follow:

<chunk_id>.<artifact_tag>.<extension>

<artifact_tag> matches the artifact directory name (with - → .):

artifact directory tag
videos/ (none — videos are the canonical source)
keyframes/uniform_5s/ kf_uni5s
keyframe-captions/…/max_imgs_16/ kf_uni5s.tSSSSSS_tEEEEEE (covered frame range)
ocr/ppocrvl15/ kf_uni5s.ocr_ppocrvl15
ocr/ppocrv5icdar/ kf_uni5s.ocr_ppocrv5icdar
ocr/pagectc/ kf_uni5s.ocr_pagectc
asr/qwen3asr1p7b/ asr_qwen3asr1p7b
asr/whisperxlargev3/ asr_whisperxlargev3
embeddings/kf_uni5s-vizemb_qwen3vlemb2b/ kf_uni5s.vizemb_qwen3vlemb2b
embeddings/kf_uni5s-vizemb_qwen3vlemb8b/ kf_uni5s.vizemb_qwen3vlemb8b
embeddings/kf_uni5s-ocr_ppocrvl15-txtemb_qwen3emb8b/ kf_uni5s.ocr_ppocrvl15.txtemb_qwen3emb8b
embeddings/kf_uni5s-ocr_ppocrv5icdar-txtemb_qwen3emb8b/ kf_uni5s.ocr_ppocrv5icdar.txtemb_qwen3emb8b
embeddings/kf_uni5s-ocr_pagectc-txtemb_qwen3emb8b/ kf_uni5s.ocr_pagectc.txtemb_qwen3emb8b

So:

videos/shard_000000.tar
    XM5xOIzL_vSkGAKR_0000.mp4
    XM5xOIzL_vSkGAKR_0000.json

keyframes/uniform_5s/shard_000000.tar
    XM5xOIzL_vSkGAKR_0000.kf_uni5s.json
    XM5xOIzL_vSkGAKR_0000.kf_uni5s.t000000.jpg
    XM5xOIzL_vSkGAKR_0000.kf_uni5s.t000005.jpg
    …

ocr/ppocrvl15/shard_000000.tar
    XM5xOIzL_vSkGAKR_0000.kf_uni5s.ocr_ppocrvl15.jsonl

embeddings/kf_uni5s-vizemb_qwen3vlemb8b/shard_000000.tar
    XM5xOIzL_vSkGAKR_0000.kf_uni5s.vizemb_qwen3vlemb8b.npz

embeddings/kf_uni5s-ocr_ppocrvl15-txtemb_qwen3emb8b/shard_000000.tar
    XM5xOIzL_vSkGAKR_0000.kf_uni5s.ocr_ppocrvl15.txtemb_qwen3emb8b.npz

If you unpack shard_000000.tar from every artifact into one directory, the files for a given chunk sort together and don't collide.

The stem before the first . is always the chunk_id — this is what WebDataset uses to group multi-artifact records for the same chunk into one sample.


Tracing IDs across artifacts

Pick any artifact and you can walk to any other.

From an OCR record → its keyframe image

The OCR jsonl line:

{"frame": "t000005", "raw": "...", "cleaned": "...", "txt": "..."}

…sits inside ocr/ppocrvl15/shard_NNN.tar under the member <chunk_id>.kf_uni5s.ocr_ppocrvl15.jsonl. The corresponding image lives at the same (shard, chunk_id) in the keyframes artifact:

keyframes/uniform_5s/shard_NNN.tar
    └── <chunk_id>.kf_uni5s.t000005.jpg     ← same `frame` value

The frame field in the OCR line and the tNNNNNN segment of the jpg filename are the same string. Trivially joinable by membership.

From a keyframe / OCR record → its source video

Locate the chunk in videos/:

videos/shard_NNN.tar
    └── <chunk_id>.mp4
    └── <chunk_id>.json

shard_NNN is the same number you read the keyframe / OCR from (every artifact shards identically). The chunk's .json gives you the offset back into the original full-length video:

{
  "video_id":         "XM5xOIzL_vSkGAKR",
  "chunk_index":      0,
  "chunk_count":      1,
  "chunk_start_sec":  0.0,
  "chunk_end_sec":    36.801,
  …
}

The original-video timestamp of frame tNNNNNN is chunk_start_sec + int(NNNNNN). For single-chunk videos chunk_start_sec == 0, so the keyframe second is the same in chunk time and source time.

From an embedding row → its keyframe and OCR record

Inside an embedding npz:

data = np.load("...npz")
data["keyframe_ids"]   # array(['t000000', 't000005', ...])
data["embeddings"]     # (N, D) float32, L2-normalised

Row i of embeddings describes frame keyframe_ids[i] of chunk_id (the .npz filename's stem). That same tNNNNNN value indexes the keyframe jpg and the OCR jsonl line for that frame.

For the text-embedding artifact: keyframe_ids may be a strict subset of the OCR frame values (frames with empty txt are dropped — see per-artifact details below).

Finding which shard holds a chunk_id

Every catalog has a shard_index column. One line:

shard = pd.read_csv("multivent-raw/videos/catalog.csv") \
          .query(f"chunk_id == '{CID}'")["shard_index"].iat[0]

The answer is the same regardless of which artifact's catalog you check.


Per-artifact details

videos/

Per chunk: one .mp4 (H.264 video, AAC audio where present) and one .json with the chunk's location in the source video.

Chunk JSON schema:

{
  "chunk_id":           "XM5xOIzL_vSkGAKR_0000",
  "video_id":           "XM5xOIzL_vSkGAKR",
  "chunk_index":        0,
  "duration_sec":       36.801,
  "source_duration_sec":36.801,
  "chunk_count":        1,
  "chunk_start_sec":    0.0,
  "chunk_end_sec":      36.801,
  "width":              360,
  "height":             640,
  "fps":                25.0
}

videos/catalog.csv: chunk_id, video_id, chunk_index, chunk_count, shard_index, duration_sec, chunk_start_sec, chunk_end_sec, size_bytes, vcodec, acodec.

keyframes/uniform_5s/

Per chunk: one .kf_uni5s.json (chunk-level metadata, extends the videos JSON with frame info) and N .kf_uni5s.tNNNNNN.jpg images, one every 5 s.

The chunk JSON adds:

"frame_count":          8,
"frame_period_sec":     5.0,
"frame_timestamps_sec": [0.0, 5.0, 10.0, 15.0, 20.0, 25.0, 30.0, 35.0]

keyframes/uniform_5s/catalog.csv: chunk_id, video_id, chunk_index, shard_index, chunk_count, frame_count, duration_sec.

ocr/ppocrvl15/

Per chunk: one .kf_uni5s.ocr_ppocrvl15.jsonl, with one JSON object per keyframe (in tNNNNNN order, length == frame_count).

Per-line schema:

{
  "frame":   "t000000",
  "raw":     "...model output with <|LOC_NNN|> coordinate tokens...",
  "cleaned": "...same content, with repetition-loop artifacts trimmed; LOC tokens preserved...",
  "txt":     "...LOC tokens stripped, whitespace tidied — ready for grep / text embedders..."
}

Most consumers want txt. Use cleaned if you need the OCR's spatial layout tokens (each <|LOC_N|> is a coordinate index 0–999). raw is preserved verbatim for anyone who wants pre-cleanup output.

ocr/ppocrvl15/catalog.csv: chunk_id, video_id, chunk_index, shard_index, n_frames, n_frames_repetition_cleaned.

asr/{qwen3asr1p7b,whisperxlargev3}/

Per chunk: one .asr_<backend>.json with language detection, voice-activity detection, speaker diarization, and a transcript. Both backends share one wrapper schema (abridged):

{
  "chunk_id":     "XM5xOIzL_vSkGAKR_0000",
  "video_id":     "XM5xOIzL_vSkGAKR",
  "chunk_index":  0,
  "duration_s":   38.824,
  "language":     {"hint": null, "detected": "Russian", "from_dir": false},
  "asr_backend":  "qwen",                                    // or "whisperx"
  "vad":          [{"start": 1.19, "end": 4.05}, ...],
  "diarization":  [{"start": 1.19, "end": 4.05, "speaker_id": "SPEAKER_00"}, ...],
  "overlap":      [...],
  "embeddings":   [{"speaker_id": "SPEAKER_00", "vector": [...], "n_turns": 3}, ...],
  "transcript": {
    "timestamp_resolution": "segment",                       // "word" for whisperx
    "segments": [{"start": 1.19, "end": 4.05, "text": "...", "speaker_id": "SPEAKER_00", "flags": []}, ...],
    "words":    []                                           // whisperx fills this with word-level timings
  },
  "qc_flags": []
}

The chunk's transcript text is the concatenation of transcript["segments"][*]["text"]. qwen3asr1p7b produces segment-level timestamps; whisperxlargev3 additionally fills transcript["words"] with word-level timings. The embeddings field holds per-speaker voice embeddings from diarization — unrelated to the embeddings/ artifact directory.

keyframe-captions/uniform_5s_qwen-multiframe-ic/max_imgs_16/

Per chunk: one or more .txt members, each an English prose caption of a window of up to 16 consecutive keyframes, produced by Qwen multi-frame image captioning. The member name encodes the covered frame range:

<chunk_id>.kf_uni5s.t000000_t000075.txt     ← caption of frames t000000 … t000075

Chunks with ≤ 16 keyframes get a single member covering the whole chunk; longer chunks get one member per 16-keyframe window. No catalog.csv for this artifact — enumerate members from the tars, or take chunk lists from videos/catalog.csv.

embeddings/kf_uni5s-vizemb_qwen3vlemb{2b,8b}/ (vision)

Per chunk: one .npz with the keyframes encoded by Qwen3-VL-Embedding (2B or 8B).

key shape dtype
keyframe_ids (N,) <U7
embeddings (N, D) float32

D = 2048 for the 2B model, D = 4096 for the 8B model. Row i describes frame keyframe_ids[i]. Embeddings are L2-normalised, so cosine similarity == inner product.

catalog.csv: chunk_id, video_id, chunk_index, shard_index, n_frames, dim.

embeddings/kf_uni5s-ocr_{ppocrvl15,ppocrv5icdar,pagectc}-txtemb_qwen3emb8b/ (text)

Per chunk: one .npz with the OCR text of each keyframe encoded by Qwen3-Embedding-8B. Three sibling artifacts, one per OCR backend.

key shape dtype
keyframe_ids (M,) <U7
embeddings (M, 4096) float32

M ≤ N. Frames whose source-OCR text was empty are skipped (no text → nothing to embed). Chunks with no text in any frame produce no member in the tar at all; their catalog row has n_frames_embedded == 0. The zero-coverage rate varies by OCR backend:

backend source field zero-emb chunks
ppocrvl15 txt ≈ 11 %
ppocrv5icdar text ≈ 8 %
pagectc text ≈ 3 %

catalog.csv (same shape for all three): chunk_id, video_id, chunk_index, shard_index, n_frames_embedded, n_frames_skipped, dim.


Catalogs

One catalog.csv per sharded artifact directory (keyframe-captions/ is the exception — it has none), 143,288 rows each. They all share the canonical prefix (chunk_id, video_id, chunk_index, shard_index), so any pair joins on chunk_id:

import pandas as pd

videos = pd.read_csv("multivent-raw/videos/catalog.csv")
ocr    = pd.read_csv("multivent-raw/ocr/ppocrvl15/catalog.csv")
txtemb = pd.read_csv("multivent-raw/embeddings/kf_uni5s-ocr_ppocrvl15-txtemb_qwen3emb8b/catalog.csv")

df = videos.merge(ocr, on="chunk_id").merge(txtemb, on="chunk_id",
                                            suffixes=("_ocr", "_txtemb"))
# guaranteed 143,288 rows

Loading examples

Plain tarfile

import tarfile, json

with tarfile.open("multivent-raw/ocr/ppocrvl15/shard_000000.tar") as tf:
    for m in tf:
        if not m.name.endswith(".jsonl"):
            continue
        for line in tf.extractfile(m).read().decode().splitlines():
            rec = json.loads(line)
            print(rec["frame"], rec["txt"][:60])
        break

One chunk's npz, by name

import io, tarfile, numpy as np

CID = "XM5xOIzL_vSkGAKR_0000"
SHARD = 0
TAR = f"multivent-raw/embeddings/kf_uni5s-vizemb_qwen3vlemb8b/shard_{SHARD:06d}.tar"
MEMBER = f"{CID}.kf_uni5s.vizemb_qwen3vlemb8b.npz"

with tarfile.open(TAR) as tf:
    blob = tf.extractfile(MEMBER).read()
data = np.load(io.BytesIO(blob))
print(data["keyframe_ids"])      # ['t000000' 't000005' ...]
print(data["embeddings"].shape)  # (N, 4096)

WebDataset (multi-artifact, joined by chunk_id)

import webdataset as wds

url = "multivent-raw/{videos,embeddings/kf_uni5s-ocr_ppocrvl15-txtemb_qwen3emb8b}/shard_000000.tar"
ds = wds.WebDataset(url, shardshuffle=False).decode()
for sample in ds:
    chunk_id = sample["__key__"]
    mp4_bytes = sample.get("mp4")
    txt_emb   = sample.get("kf_uni5s.ocr_ppocrvl15.txtemb_qwen3emb8b.npz")
    ...

Sharding

667 shards of ~210 chunks each (range 121–272). A chunk lives in exactly one shard across all artifacts: shard 42 of videos/ and shard 42 of any other artifact describe the same set of chunks. Shards are independent, so you can shuffle shard order then iterate within for IID-ish batches. wds.WebDataset(shardshuffle=True) does this automatically.

Downloads last month
34