Dataset Viewer
Auto-converted to Parquet Duplicate
Search is not available for this dataset
image
imagewidth (px)
1.28k
1.92k
label
class label
18 classes
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
0video_10717
End of preview. Expand in Data Studio

VLM point-tracking benchmark

The current Table 1 input release is benchmark_312_chen_v1: 312 clips, 14,896 two-image point-correspondence requests, 17,427 unique input images. It freezes the red reference marker, embedded zoom inset, target frame, rendered Chen-v1 prompt, ground truth, motion labels and category mapping.

Download instructions and exact configuration status are in the release README. No model inference was performed to prepare this release. GPT default reasoning and Muse minimal are configured and locally checked. Gemini/Claude adapters and the proposed gateway are still pending; this is not a completed four-model result.

Legacy artifacts

The top-level queries/strat1200_red/ artifacts are the earlier 300-clip grouped query release. They are preserved for provenance and are not the current main table inputs. Their legacy STIR point identities were unverified; the new release excludes STIR entirely and adds 20 TAP-Vid Kubric clips. Use the new release for fresh evaluation instead of mixing these query/GT files.

This collection aggregates source datasets with their original terms and attributions; those terms still apply. No new blanket license is granted here.

Download and run

Verified input/runner revision: ed6a85e7bcac334cdba75aa25ee33a989d6eb1e4. Download only this release (about 4.0 GB), then verify and unpack it:

hf download 2inf/tracking_bench_vlm --repo-type dataset \
  --revision ed6a85e7bcac334cdba75aa25ee33a989d6eb1e4 \
  --include 'releases/benchmark_312_chen_v1/*' --local-dir hf_download
(cd hf_download/releases/benchmark_312_chen_v1 && sha256sum -c SHA256SUMS)
mkdir -p benchmark_312_run
tar -xf hf_download/releases/benchmark_312_chen_v1/runner.tar -C benchmark_312_run
tar -xf hf_download/releases/benchmark_312_chen_v1/inputs.tar -C benchmark_312_run
cd benchmark_312_run
python3.11 -m venv .venv
.venv/bin/pip install -r requirements-benchmark.txt
.venv/bin/python scripts/launcher.py configs/experiments/benchmark_312_chen_gpt6.yaml --dry-run --limit 1
.venv/bin/python scripts/launcher.py configs/experiments/benchmark_312_chen_muse.yaml --dry-run --limit 1

Install huggingface_hub first if the hf download command is unavailable. No GPU is needed. Supply API_KEY for GPT and MUSE_API_KEY for Muse through private environment variables; do not put credentials in the shared run artifacts.

After checking the dry-run output, these are the inference commands for the person launching the benchmark (they were not executed during preparation):

.venv/bin/python scripts/launcher.py configs/experiments/benchmark_312_chen_gpt6.yaml --run-id chen_v1
.venv/bin/python scripts/launcher.py configs/experiments/benchmark_312_chen_muse.yaml --run-id chen_v1

The named run resumes previously answered requests. Keep the full output folders under experiments/benchmark_312_chen_gpt6/chen_v1/ and experiments/benchmark_312_chen_muse/chen_v1/. The raw call journals and append-only attempt log are part of the result, not disposable debugging output.

Downloads last month
243