The ToolRL model trained for tool use through GRPO
Cheng Qian
chengq9
AI & ML interests
Agent, Tool Learning
Recent Activity
upvoted a paper 2 days ago
Rethinking On-Policy Distillation of Large Language Models II: One Training Example authored a paper 22 days ago
AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses authored a paper 22 days ago
ACCORD: Action-Conditioned Contextual Grounding for Language Agents