WebWorld: The Browser as a World Model for Self-Improving Web Code Paper • 2608.30530 • Published 11 days ago • 10
EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning Paper • 2608.06197 • Published Aug 6 • 47
Self-State Attacks on Self-Hosted AI Agents: How Far Can OS Defenses Go? Paper • 2607.17986 • Published Jul 20 • 6
LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget Paper • 2607.14952 • Published Jul 16 • 213
WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces Paper • 2606.09426 • Published Jun 8 • 108
From Model Scaling to System Scaling: Scaling the Harness in Agentic AI Paper • 2605.26112 • Published May 25 • 10
DCAgent3/medagentbench_g1_diverse_tezos_top4_3160_8b_20260602_100923 Viewer • Updated Jun 2 • 1.49k • 13 • 1