DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments? Paper • 2608.10366 • Published 4 days ago • 9
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Paper • 2608.05747 • Published 9 days ago • 46
MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations Paper • 2607.28956 • Published 15 days ago • 96
GradCuit: Credit-Assigned Gradient Flow Enables Robust and Interpretable Test-Time Latent Reasoning Paper • 2608.02585 • Published 12 days ago • 24
Scaling Properties of Text Conditioning in Visual Generation Paper • 2607.29679 • Published 15 days ago • 38