Datasets:
CoinVE-Bench
✨ Highlights
- Compositional Editing Benchmark: CoinVE-Bench is a benchmark dedicated to compositional (multi-instruction) video editing evaluation, featuring 361 test cases with 2–5 instructions per sample.
- Region-Aware Evaluation: Per-instruction masks enable fine-grained checks of whether each edit is confined to the correct region.
- Multi-Dimensional Metrics: Comprehensive evaluation across editing accuracy, physical naturalness, and semantic preservation.
- VLLM-Based Scoring: Leveraging vision-language models (e.g., Gemini) for human-aligned assessment.
📊 Benchmark Statistics
| Metric | Value |
|---|---|
| Total test samples | 361 |
| Instructions per sample | 2–5 |
| Checklist questions | 4,131 |
| Evaluation dimensions | 3 (Editing Accuracy / Physical Naturalness / Semantic Preservation) |
Scoring Dimensions & Metrics
All metrics are scored by Gemini 3.6 Flash (default; configurable via --model) using a checklist-based VLLM protocol.
| Dimension | Metric (Abbr.) | Range | # Questions | Description |
|---|---|---|---|---|
| Editing Accuracy (per instruction) | Semantic Accuracy (SA) | [0, 100] | 1,081 | Correct execution of the intended edit semantics. |
| Scope Accuracy (SPA) | [0, 100] | 379 | Correct edit localization without leakage or interference. | |
| Editing Persistence (EP) | [0, 100] | 923 | Temporal consistency of the edit throughout the video. | |
| Physical Naturalness | Appearance Naturalness (AN) | [0, 100] | 356 | Natural blending of lighting, shadows, textures, and style. |
| Scale Consistency (SC) | [0, 100] | 521 | Plausibility of the edited object's scale and perspective. | |
| Motion Naturalness (MN) | [0, 100] | 447 | Plausibility of motion and physical interactions. | |
| Semantic Preservation | Content Preservation (CP) | [0, 100] | 424 | Preservation of non-edited regions, objects, and structures. |
Scoring formulas:
- ACC-type metrics (SA / SPA / EP / AN / SC / MN):
score = num_correct / num_questions × 100 - Score-type metric (CP):
score = mean(score) × 10(each question is scored 0–10, then scaled to [0, 100]) - Dimension score = arithmetic mean of its metric
score_100values
Question Types
| Type | # Questions | Videos Uploaded | Answer Format |
|---|---|---|---|
| Single True/False (Single-TF) | 2,524 | Edited Video only | Yes / No |
| A/B Multiple Choice Question (AB-MCQ) | 609 | Edited Video only | A / B / A and B |
| Dual True/False (Dual-TF) | 574 | Source Video + Edited Video | Yes / No |
| Score Multiple Choice Question (Score-MCQ) | 424 | Source Video + Edited Video | 0–10 score |
📁 Dataset Structure
CoinVE-Bench/
├── src_videos/
│ ├── 9CicJDFN1TA_9_0to183.mp4
│ ├── gWG0LiuZnFQ_8_0to137.mp4
│ ├── JwVDrRmyoXc_51_0to172.mp4
│ └── ...
└── checklist_json/
└── coinve-bench-361-checklist.json
Source videos are named <youtube_id>_<clip_index>_<start>to<end>.mp4 and are referenced by the src_video field of each case in the checklist (e.g. src_videos/9CicJDFN1TA_9_0to183.mp4).
🔧 Evaluation
Evaluation is a two-stage pipeline: (1) run your editing model on the 361 source videos to produce edited videos, (2) score them with the Gemini-based evaluator.
1. Run Inference on CoinVE-Bench
Generate one edited video per test case using your model. For CoinVE-Edit, use the inference scripts in the CoinVE-Edit model repository.
2. Prepare Edited Videos
The evaluator expects one flat directory of edited MP4 files with a strict naming convention:
edited_videos/your_model/
├── 0.mp4
├── 1.mp4
├── 2.mp4
└── ...
└── 360.mp4
- Format: MP4 (
.mp4). - Naming:
<case_id>.mp4where<case_id>is the integeridfield of each case inchecklist_json/coinve-bench-361-checklist.json(range 0–360, no zero-padding). - Source video pairing: Video A (original) is read from
case["src_video"]inside the checklist; you do not need to place source videos in this directory. - Missing videos: if a case's edited video is absent, all its questions are recorded with
model_answer = nulland counted as failures in the final score.
Note for CoinVE-Edit users:
infer_coinve_bench.pysaves a standalone edited video attgt_videos/<case_id>.mp4per case, already in the{id}.mp4naming convention expected by the evaluator. You can pass thetgt_videos/directory directly as--gen-video-dir.
3. Run Gemini Evaluation
The evaluation code is provided in the CoinVE-Edit repository. Set your Gemini API key and run:
export GEMINI_API_KEY=your_key_here
python eval_coinbench_gemini_public.py \
--input ./checklist_json/coinve-bench-361-checklist.json \
--gen-video-dir ./edited_videos/your_model \
--prompt-dir ./system_prompts \
--out-dir ./results/your_model \
--workers 8 \
--model gemini-3.6-flash \
--timestamp
The script performs everything end-to-end — Gemini calls, per-question scoring, dimension/metric aggregation, and report writing. No separate aggregation step is needed.
📈 Performance Comparisons
Metric groups: SA / SPA / EP → Edit. Acc. · AN / SC / MN → Phys. Natural. · CP → Seman. Pres.
| Model | SA | SPA | EP | AN | SC | MN | CP |
|---|---|---|---|---|---|---|---|
| Seedance 2.0 | 85.34 | 87.71 | 88.08 | 93.19 | 95.84 | 92.87 | 93.91 |
| Kling O3 | 86.91 | 80.93 | 89.06 | 92.55 | 90.30 | 93.91 | 84.51 |
| VACE | 3.98 | 17.15 | 6.50 | 26.69 | 13.82 | 15.21 | 87.83 |
| Ditto | 34.69 | 36.41 | 40.85 | 35.96 | 47.79 | 38.48 | 51.98 |
| VINO | 83.63 | 66.75 | 89.06 | 78.09 | 82.34 | 85.91 | 61.70 |
| OmniWeaving | 59.67 | 55.94 | 61.11 | 54.49 | 66.03 | 65.10 | 75.09 |
| KiWiEdit | 76.50 | 69.92 | 80.28 | 78.37 | 78.50 | 80.76 | 70.31 |
| SAMA | 75.58 | 73.35 | 79.63 | 83.43 | 83.88 | 88.14 | 90.08 |
| CoinVE-Edit | 87.97 | 89.45 | 89.60 | 91.85 | 91.17 | 95.30 | 90.83 |
📜 Citation
If you find CoinVE-Bench useful for your research, please cite our work:
@article{coinve200k,
title={CoinVE-200K: A Large-Scale High-Quality Dataset for Compositional Instruction-Guided Video Editing},
author={Long, Fuchen and Wang, Cong and Gao, Zitao and Zhong, Wenhao and Cheng, Yu and Hou, Xiaolu and Li, Yan and Cao, Xiao and Sun, Xinlong and Chen, Xi and Liu, Yu},
journal={arXiv preprint arXiv:2608.17566},
year={2026}
}
✉️ Contact
For any questions, issues, or collaborations, please feel free to contact [email protected].
💖 Acknowledgement
Our benchmark construction and evaluation protocol are inspired by ReCo-Bench and OpenVE-Bench. Thanks to the contributors of all these remarkable projects!
- Downloads last month
- 7