Memorizon: Training World Models Beyond Their Context Window
Paper โข 2610.00544 โข Published โข 10
Tingting Liao ยท
Xuezhi Liang ยท
Hao Li ยท
Guangyi Liu
Institute of Foundation Models (IFM), MBZUAI
| Model | Memorizon, 4-step distilled (Self-Forcing + DMD, CFG distilled in) |
| Base | Wan2.2-TI2V-5B |
| Input | one image + keyboard actions or a camera trajectory |
| Output | 864ร480, 16 fps, generated one second at a time |
| Sampling | 4 steps, no CFG (read from the config) |
git clone https://github.com/TingtingLiao/memorizon.git && cd memorizon && pip install -e .
python scripts/generate.py --checkpoint Luffuly/memorizon --image photo.jpg \
--actions "w*16 l*24 w*16 l*24" --output out.mp4
from memorizon import MemorizonPipeline, actions_to_c2w, save_video
pipe = MemorizonPipeline.from_pretrained("Luffuly/memorizon")
video = pipe("photo.jpg", actions_to_c2w("w*16 l*24 w*16 l*24"))
save_video(video, "out.mp4")
| Key | Action | Key | Action |
|---|---|---|---|
w s |
forward / backward ยท 0.25 m | j l |
turn left / right ยท 7.5ยฐ |
a d |
left / right ยท 0.25 m | i k |
look up / down ยท 7.5ยฐ |
. |
stay | *N |
repeat N times |
Without a prompt, the image is captioned by Qwen3-VL-8B in the training format
(memorizon/caption.py). Custom prompts should follow it โ a perspective prefix,
then one paragraph:
First-person perspective โ character not visible. A winding paved path curves beneath a canopy of vibrant pink cherry blossoms, โฆ
Instead of actions, pass --trajectory a [T, 4, 4] array of camera-to-world poses,
T = 1 + 4 ร seconds, in the first frame's coordinates (OpenCV axes), translations
in metres / 4.
@article{memorizon2026,
title = {Memorizon: Training World Models Beyond Their Context Window},
author = {Tingting Liao, Xuezhi Liang, Hao Li, Guangyi Liu},
journal = {arXiv preprint arXiv:2610.00544},
year = {2026}
}
Base model
Wan-AI/Wan2.2-TI2V-5B-Diffusers