Demystifying Agent Skills: Why They Work-Until They Don't Paper • 2608.14036 • Published 10 days ago • 162
StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling Paper • 2608.15089 • Published 9 days ago • 436
Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development Paper • 2608.13417 • Published 11 days ago • 54
Alaya-EVOKE: From Linear-Scaling Supervision to Endless World Paper • 2608.13546 • Published 11 days ago • 132
AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses Paper • 2608.12307 • Published 12 days ago • 113
SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs Paper • 2608.03573 • Published 18 days ago • 57
Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence Paper • 2608.10720 • Published 13 days ago • 16
Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA Paper • 2608.09819 • Published 14 days ago • 338
OneEmo: A Unified Multimodal Reasoning Model for Emotion Perception, Understanding, and Interaction Paper • 2608.06013 • Published 18 days ago • 5
SimWAM: A Simple World Action Model for End-to-End Autonomous Driving Paper • 2608.07468 • Published 17 days ago • 106
SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding Paper • 2608.05137 • Published 14 days ago • 27
DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces Paper • 2608.03451 • Published 20 days ago • 33
When Teachers Mislead: Spurious-Signal-Aware On-Policy Distillation Paper • 2608.03632 • Published 20 days ago • 23
Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing Paper • 2608.02711 • Published 20 days ago • 91
CURV: Enhancing Chart Understanding Through Curriculum Visual Grounded Reasoning Paper • 2608.02833 • Published 21 days ago • 4
ST-WAM: Semantic-Temporal World Action Model for Robust Manipulation under Visual Distribution Shifts Paper • 2607.28993 • Published 24 days ago • 7
LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks Paper • 2608.01964 • Published 21 days ago • 180
ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction Paper • 2607.29677 • Published 24 days ago • 24
LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger Paper • 2607.28374 • Published 25 days ago • 15