Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds Paper • 2608.23383 • Published 6 days ago • 19
EchoWM: Open and Enterable Omnimodal World Models Paper • 2608.23189 • Published 6 days ago • 76
SANA-WM Collection SANA-WM: Efficient Minute-Scale World Modeling with Hybrid Linear Diffusion Transformer • 6 items • Updated Jun 11 • 7
Self Gradient Forcing: Native Long Video Extrapolation Paper • 2607.20368 • Published Jul 22 • 36
Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models Paper • 2606.25041 • Published Jun 23 • 122
JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence Paper • 2606.14777 • Published Jun 10 • 217
Echo-Memory: A Controlled Study of Memory in Action World Models Paper • 2606.09803 • Published Jun 8 • 33
Echo-Infinity: Learning Evolving Memory for Real-Time Infinite Video Generation Paper • 2606.04527 • Published Jun 3 • 28
Learning A Unified Risk Map for Autonomous Driving in Partially Observable Environments Paper • 2605.22189 • Published May 21 • 8
Learning A Unified Risk Map for Autonomous Driving in Partially Observable Environments Paper • 2605.22189 • Published May 21 • 8
view article Article Introducing NVIDIA Nemotron 3 Nano Omni: Long-Context Multimodal Intelligence for Documents, Audio and Video Agents nvidia • Apr 28 • 64