GraphSkillEvo: Evolutionary Optimization of Graph-Structured Agent Skills Paper • 2609.21749 • Published 15 days ago • 19
GraphSkillEvo: Evolutionary Optimization of Graph-Structured Agent Skills Paper • 2609.21749 • Published 15 days ago • 19
Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO Paper • 2608.27351 • Published Aug 27 • 22
Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO Paper • 2608.27351 • Published Aug 27 • 22
Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements Paper • 2608.17310 • Published Aug 18 • 110
Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements Paper • 2608.17310 • Published Aug 18 • 110
One Token per Multimodal Evidence: Latent Memory for Resource-Constrained QA Paper • 2606.10572 • Published Jun 9 • 17
One Token per Multimodal Evidence: Latent Memory for Resource-Constrained QA Paper • 2606.10572 • Published Jun 9 • 17
Beyond Imitation: Reinforcement Learning for Active Latent Planning Paper • 2601.21598 • Published Jan 29 • 10
Beyond Imitation: Reinforcement Learning for Active Latent Planning Paper • 2601.21598 • Published Jan 29 • 10
SofT-GRPO: Surpassing Discrete-Token LLM Reinforcement Learning via Gumbel-Reparameterized Soft-Thinking Policy Optimization Paper • 2511.06411 • Published Nov 9, 2025 • 19
Monte Carlo Tree Search for Comprehensive Exploration in LLM-Based Automatic Heuristic Design Paper • 2501.08603 • Published Jan 15, 2025
UDC: A Unified Neural Divide-and-Conquer Framework for Large-Scale Combinatorial Optimization Problems Paper • 2407.00312 • Published Jun 29, 2024
Reasoning-CV: Fine-tuning Powerful Reasoning LLMs for Knowledge-Assisted Claim Verification Paper • 2505.12348 • Published May 18, 2025