NeMo-DCR: Bit-Exact Delta-Compressed Refit for Scalable Agentic RL at Trillion-Parameter Scale Paper • 2610.08430 • Published 5 days ago • 24
Retrieval-Augmented Skill Optimization via Cross-Harness Adaptation Paper • 2609.38024 • Published 12 days ago • 64
The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks Paper • 2609.25804 • Published 19 days ago • 164
onPanda: Efficient Annotation of On-Policy Alignment Data for LLMs and Agents via Token-Level Correction Paper • 2609.24983 • Published 20 days ago • 55
IntBMoE: Integrating Block-Level Conditioning into Expert Composition for Full-Participation Mixture-of-Experts Paper • 2609.21346 • Published 23 days ago • 101
Calibrating Teacher--Student Discrepancy for On-Policy Distillation Paper • 2609.21619 • Published 23 days ago • 13
UFO: Chain-of-Evaluation for Omni-Condition Alignment in Multi-Modal Image Generation Paper • 2609.12397 • Published 24 days ago • 46
ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement Paper • 2609.14857 • Published 27 days ago • 215
Dream-RSI: Recursive Self-Improvement through Evolving Worlds Paper • 2609.14858 • Published 27 days ago • 252
Discovery Foundation Models: Toward Open-Ended Discovery Intelligence Paper • 2609.15973 • Published 27 days ago • 33
An Empirical Study on Zero-Data Bootstrapping for Conversational Recommender Systems Paper • 2504.15476 • Published Aug 28 • 2
Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation Paper • 2609.08084 • Published Sep 8 • 72
NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness Paper • 2609.08183 • Published Sep 8 • 321
Bilevel Coordinated Reflection: A Game-Theoretic Approach to Multi-Agent LLM Systems Paper • 2609.02750 • Published Sep 2 • 145
Rethinking On-Policy Distillation of Large Language Models II: One Training Example Paper • 2609.04172 • Published Sep 3 • 104