Abstract
EchoWM is an omnimodal world model that generates synchronized high-resolution video, sound, music, and speech while following continuous 6-DoF navigation trajectories across first- and third-person views.
We present EchoWM, an omnimodal world model for enterable generative media that responds to continuous navigation while jointly generating 720p video, environmental sound, music and speech. We organize interaction around camera intent: in first-person scenes, it specifies observer motion, while in third-person scenes, camera--character dynamics are learned from data without view-specific controllers. Discrete commands and continuous poses are mapped to a shared metric-scale relative 6-DoF trajectory, with dataset-level calibration preserving motion magnitude across heterogeneous data. To jointly learn audio-visual generation and trajectory control, we construct a complementary data engine and adopt progressive training followed by autoregressive post-training for long-horizon generation. Extensive evaluations show that \model achieves strong trajectory following and high visual quality on public world-model benchmarks, supporting both first- and third-person interaction across varied subjects, and maintaining synchronized environmental sound and speech over long-horizon generation.
Community
An omnimodal world model for generative media that responds to continuous navigation while video, environmental sound, music, and speech evolve together.
insane work
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report (2026)
- ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU (2026)
- Vorch-Omni: Multi-Task Orchestration of Sight and Sound (2026)
- AlayaWorld: Long-Horizon and Playable Video World Generation (2026)
- Vorch-Streamer: Extending Human Audio-Visual Generation to Real-Time Long-Form Streaming (2026)
- Wonder: Video World Model Done Better (2026)
- MiniWorld: Democratizing the Training of Video World Models from Scratch (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2608.23189 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper