Knowing When to Quit: Diagnosing and Training LLMs to Abort Futile Reasoning Paper • 2607.29211 • Published 17 days ago • 11
DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments? Paper • 2608.10366 • Published 6 days ago • 10
HarnessOpt-Bench: Evaluating LLMs at Harness Optimization Paper • 2608.06301 • Published 11 days ago • 34
Any-OPD: Heterogeneous On-Policy Distillation for Flow-Matching Models via Representation-Space Bridging Paper • 2608.03316 • Published 13 days ago • 26
ethanfel/Qwen3-VL-32B-Ultra-Heretic-H3-ComfyUI-INT8-ConvRot Image-Text-to-Text • Updated 11 days ago • 501