SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research? Paper • 2609.09113 • Published 18 days ago • 20
SWE-Touch: Benchmarking Coding Agents When Users Touch the Code Paper • 2608.02499 • Published Aug 3 • 25
Bottom-up Policy Optimization: Your Language Model Policy Secretly Contains Internal Policies Paper • 2512.19673 • Published Dec 22, 2025 • 66