Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation Paper • 2607.27372 • Published 5 days ago • 13
AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis Paper • 2607.28618 • Published 4 days ago • 292
Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning Paper • 2607.21653 • Published 12 days ago • 31
Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills Paper • 2607.22529 • Published 10 days ago • 47
Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction Paper • 2607.20911 • Published 11 days ago • 26
DSWorld: A Data Science World Model for Efficient Autonomous Agents Paper • 2607.15901 • Published 17 days ago • 12
SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Paper • 2607.14777 • Published 18 days ago • 103
DrugGen 2: A disease-aware language model for enhancing drug discovery Paper • 2607.08404 • Published 25 days ago • 20
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Paper • 2607.04412 • Published 29 days ago • 35
RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies Paper • 2607.04434 • Published 27 days ago • 15
SkillCoach: Self-Evolving Rubrics for Evaluating and Enhancing Agentic Skill-Use Paper • 2607.01874 • Published Jul 2 • 23
AgenticDataBench: A Comprehensive Benchmark for Data Agents Paper • 2607.01647 • Published Jul 2 • 37
AutoTrainess: Teaching Language Models to Improve Language Models Autonomously Paper • 2606.31551 • Published Jun 30 • 24
Autonomous Scientific Discovery via Iterative Meta-Reflection Paper • 2607.01131 • Published Jul 1 • 8
Evolution Fine-Tuning Collection Internalizing Discovery Capability into LLM • 10 items • Updated Jul 1 • 1
Evolution Fine-Tuning: Learning to Discover Across 371 Optimization Tasks Paper • 2606.29082 • Published Jun 27 • 42
Agentic Abstention: Do Agents Know When to Stop Instead of Act? Paper • 2606.28733 • Published Jun 27 • 150
OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks Paper • 2606.29537 • Published Jun 28 • 24