CoinVE-200K: A Large-Scale High-Quality Dataset for Compositional Instruction-Guided Video Editing
Abstract
A new dataset, benchmark, and 22B model enable compositional instruction-guided video editing with multi-region attention and temporal coherence.
The quality and diversity of instruction-based video editing datasets are steadily improving, yet existing datasets mainly focus on single editing operations and fall short in supporting compositional instruction-guided video editing. In particular, multiple editing intents must be jointly understood and faithfully executed within the same video. To address this issue, we introduce CoinVE-200K, a large-scale, high-quality dataset for Compositional Instruction-Guided Video Editing. CoinVE-200K contains 1080p video-editing pairs of up to 201 frames, covering diverse compositional scenarios where each sample involves 2 to 5 atomic editing operations. The instructions target humans, objects, and backgrounds, and cover edit types such as addition, removal, modification, and stylization. All samples are built through a carefully designed generation and filtering pipeline to ensure instruction faithfulness, visual quality, temporal consistency, and compositional diversity. We also introduce CoinVE-Bench, a benchmark for compositional-instruction video editing across diverse subjects, operation types, and instruction complexities. Furthermore, we present CoinVE-Edit, a 22B compositional video editing model built upon Wan2.1-T2V-14B and Qwen3-VL-8B-Instruct. CoinVE-Edit disentangles region-aware attention for different editing instructions, enabling precise multi-region editing while preserving irrelevant content and temporal coherence. Experiments on CoinVE-Bench show that CoinVE-Edit achieves strong performance in instruction following, compositional editing accuracy, visual quality, and temporal consistency.
Community
CoinVE-200K: A Large-Scale High-Quality Dataset for Compositional Instruction-Guided Video Editing
Project Page: https://coinve200k.github.io/
Code: https://github.com/coinve200k/CoinVE-200K
Dataset: https://huggingface.co/datasets/FireCRT/CoinVE-200K
Model: https://huggingface.co/FireCRT/CoinVE-Edit
Bench: https://huggingface.co/datasets/FireCRT/CoinVE-Bench
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Goku: A Million-Scale Universal Dataset and Benchmark for Instruction-Based Video Editing (2026)
- VicEdit: Learning to Edit Videos from Visual In-Context Examples (2026)
- InstructVVT: Instruction-Driven Video Virtual Try-On without Auxiliary Spatial Priors (2026)
- CoT-Edit: Let CoT Guide Instruction Video Editing (2026)
- OmniEdit-Bench: A Comprehensive Benchmark for Instruction-based Video Editing (2026)
- CineWeaver: Training-Free Reference-Controllable Multi-Shot Long Video Generation for Cinematic Storytelling (2026)
- Vera: A Layered Diffusion Model for Content-Preserving Video Editing (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2608.17566 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 1
Datasets citing this paper 2
FireCRT/CoinVE-200K
FireCRT/CoinVE-Bench
Spaces citing this paper 1
Collections including this paper 0
No Collection including this paper