DriveZero: End-to-End Driving Beyond Human Demonstrations Paper • 2609.06055 • Published 30 days ago • 57
Lucida: Parse, Generate, and Place for Composable Real-to-Sim Scene Modeling Paper • 2608.30821 • Published Aug 31 • 69
Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios Paper • 2608.25529 • Published Aug 26 • 17
WBench: A Comprehensive Multi-turn Benchmark for Interactive Video World Model Evaluation Paper • 2605.25874 • Published May 25 • 82
Uni-Edit: Intelligent Editing Is A General Task For Unified Model Tuning Paper • 2605.21487 • Published May 20 • 21
HiAR: Efficient Autoregressive Long Video Generation via Hierarchical Denoising Paper • 2603.08703 • Published Mar 9 • 32
U4D: Uncertainty-Aware 4D World Modeling from LiDAR Sequences Paper • 2512.02982 • Published Dec 2, 2025 • 3
LongVie 2: Multimodal Controllable Ultra-Long Video World Model Paper • 2512.13604 • Published Dec 15, 2025 • 76
Architecture Decoupling Is Not All You Need For Unified Multimodal Model Paper • 2511.22663 • Published Nov 27, 2025 • 29
MCPMark: A Benchmark for Stress-Testing Realistic and Comprehensive MCP Use Paper • 2509.24002 • Published Sep 28, 2025 • 153
LongVie: Multimodal-Guided Controllable Ultra-Long Video Generation Paper • 2508.03694 • Published Aug 5, 2025 • 53