TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training Paper • 2607.05804 • Published Jul 7 • 20
AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Paper • 2608.05987 • Published 19 days ago • 100
Knowing When to Quit: Diagnosing and Training LLMs to Abort Futile Reasoning Paper • 2607.29211 • Published 25 days ago • 14
Knowing When to Quit: Diagnosing and Training LLMs to Abort Futile Reasoning Paper • 2607.29211 • Published 25 days ago • 14
Improving Multi-step RAG with Hypergraph-based Memory for Long-Context Complex Relational Modeling Paper • 2512.23959 • Published Dec 30, 2025 • 112
BARREL: Boundary-Aware Reasoning for Factual and Reliable LRMs Paper • 2505.13529 • Published May 18, 2025 • 12
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Paper • 2501.01830 • Published Jan 3, 2025 • 17
DeepRAG: Thinking to Retrieval Step by Step for Large Language Models Paper • 2502.01142 • Published Feb 3, 2025 • 25
DeepRAG: Thinking to Retrieval Step by Step for Large Language Models Paper • 2502.01142 • Published Feb 3, 2025 • 25
DeepRAG: Thinking to Retrieval Step by Step for Large Language Models Paper • 2502.01142 • Published Feb 3, 2025 • 25 • 2