BUT-FIT/diarizen-wavlm-large-s80-md-v2 Voice Activity Detection • Updated about 1 hour ago • 819 • 21
view article Article Unified Models for Image Understanding and Generation: Understanding Cutting-Edge Model Architectures exploding-gradients • Sep 15, 2025 • 5
DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment Paper • 2607.07820 • Published Jul 8 • 94
AnyTalker: Scaling Multi-Person Talking Video Generation with Interactivity Refinement Paper • 2511.23475 • Published Nov 28, 2025 • 43
GenCompositor: Generative Video Compositing with Diffusion Transformer Paper • 2509.02460 • Published Sep 2, 2025 • 26
Running on Zero Agents 217 IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System 🎙 217 Generate speech from text using a reference audio
Running on Zero Agents Featured 1.45k EasyControl Ghibli 🦀 1.45k New Ghibli EasyControl model is now released!!
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Paper • 2505.22647 • Published May 28, 2025 • 3