Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation Paper • 2609.08084 • Published 14 days ago • 70
TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming Paper • 2608.20958 • Published Aug 21 • 58
Zetta ζ: An Efficient Closed-Loop Embodied Harness for Self-Evolving Physical Intelligence Paper • 2608.16590 • Published Aug 17 • 151
Distill Where You Fail: Recovering Learning Signals of Negative RL-Groups from Adaptive Teacher Guidance Paper • 2608.00782 • Published Aug 1 • 18
FocusMem: Factorizing Content, Readout, and Trust in Latent GUI Memory Paper • 2608.04530 • Published Aug 5 • 14
HelloWorld: Enabling Socially Interactive Characters in Video World Models Paper • 2608.05070 • Published Aug 5 • 41
WorldCycle: Self-Verifiable Reinforcement Learning for Long-Horizon Video World Models Paper • 2608.04964 • Published Aug 5 • 14
OPD-V: Visual On-Policy Self-Distillation with Modality Balance Paper • 2608.05131 • Published Aug 6 • 15