SemanTok: Predictable Semantic Tokens for Efficient Autoregressive Video Generation
Abstract
Recent video-based world models pair the scalability of autoregressive (AR) prediction with the visual quality of diffusion models. The choice of scene tokenizer is paramount for the optimal performance of each of these, both in terms of fidelity and semantics. Flexible-length, coarse-to-fine tokenizers yield exactly that: the first coarse tokens carry the clip's global semantics while later tokens further specify details. Existing flexible tokenizers only apply a representation-alignment (REPA) loss on early decoder hidden states, a target the decoder can partly meet from its noised input instead. We introduce SemanTok, a flexible video tokenizer that feeds frozen DINO features into its encoder and adds lightweight heads that reconstruct them from each retained token prefix alone. SemanTok achieves high semantic alignment and video fidelity at every AR model size: a 201M SemanTok AR model matches or beats a VideoFlexTok AR model 3.4times its size, and larger SemanTok AR models further improve fidelity. It keeps semantic alignment on out-of-distribution classes and gives the decoder higher semantic alignment at every noise level, including pure noise. It performs well in both reconstruction and generation, and its short token prefixes are cheaper to predict and give better generation fidelity, with pixel detail deferred to later tokens.
Community
Flexible video tokenizers (e.g., VideoFlexTok) let an autoregressive (AR) model stop after any number of tokens, which condition a diffusion decoder, so the first tokens should already capture what the clip shows. SemanTok supervises this explicitly: every nested token prefix is trained to carry the clip's semantics. The resulting prefixes are cheaper to predict and lead to better generation fidelity and higher semantic alignment: a 201M SemanTok AR model matches or beats a VideoFlexTok AR model 3.4× its size.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Keep-or-Drop? Adaptive Tokenizer for Compact Video Representation (2026)
- V-RAE: Rethinking Video Latent Spaces for Generation (2026)
- W2Rep: Learning Visual Representations by Watching the World Change (2026)
- TT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining (2026)
- RESUME: Recurrent State Updates from Motion and Residual Signals for Efficient Video Language Modeling (2026)
- SeRV: Semantic-Aligned Residual Vector Quantization for American Sign Language Generation (2026)
- SemBridge: Semantic Token Anchoring for Continuous-Latent Autoregressive Speech Generation (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2610.00686 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper