GenFirst: Generation Before Reconstruction for Stable End-to-End Latent Generative Modeling Paper • 2608.29335 • Published 9 days ago • 66
Poplar: A Scalable Pipeline for Human-Centric Image Dataset Synthesis Paper • 2608.00440 • Published Aug 1 • 10
dRAE: Representation Autoencoder with Hyper-Spherical Codes Paper • 2607.22148 • Published Jul 24 • 12
SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation Paper • 2607.21553 • Published Jul 23 • 39
Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning Paper • 2607.21653 • Published Jul 22 • 32
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Paper • 2607.14935 • Published Jul 16 • 172
PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space Paper • 2607.05373 • Published Jul 6 • 66
From SRA to Self-Flow: Data Augmentation or Self-Supervision? Paper • 2607.02508 • Published Jul 2 • 15
Dockerless: Environment-Free Program Verifier for Coding Agents Paper • 2606.28436 • Published Jun 26 • 117
MIMFlow: Integrating Masked Image Modeling with Normalizing Flows for End-to-End Image Generation Paper • 2606.26016 • Published Jun 24 • 10
DiffusionBench: On Holistic Evaluation of Diffusion Transformers Paper • 2606.24888 • Published Jun 23 • 12
UniDDT: Unifying Multimodal Understanding and Generation with Decoupled Diffusion Transformer Paper • 2606.16255 • Published Jun 15 • 16