Demystifying Agent Skills: Why They Work-Until They Don't Paper • 2608.14036 • Published 6 days ago • 148
Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill Paper • 2608.11924 • Published 8 days ago • 282
PrivacyPeek: Auditing What LLM-Based Agents Acquire, Not Just What They Say Paper • 2606.00152 • Published 14 days ago • 4
Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs Paper • 2608.01755 • Published 17 days ago • 141
Can Dialects Be Steered Like Languages? Sparse Neurons and Distributed Directions in Arabic LLMs Paper • 2607.03936 • Published Jul 4 • 5
MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Paper • 2606.00793 • Published Jun 8 • 11
Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models Paper • 2606.11324 • Published Jun 9 • 172
Perception or Prejudice: Can MLLMs Go Beyond First Impressions of Personality? Paper • 2605.22109 • Published May 21 • 171
IndusAgent: Reinforcing Open-Vocabulary Industrial Anomaly Detection with Agentic Tools Paper • 2605.20682 • Published May 20 • 86
Zero-Shot Sim-to-Real Robot Learning: A Dexterous Manipulation Study on Reactive Catching Paper • 2605.09789 • Published May 10 • 6
CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence Paper • 2605.12882 • Published May 13 • 274
Preferences of a Voice-First Nation: Large-Scale Pairwise Evaluation and Preference Analysis for TTS in Indian Languages Paper • 2604.21481 • Published Apr 23 • 3
Adam's Law: Textual Frequency Law on Large Language Models Paper • 2604.02176 • Published Apr 2 • 511
Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability Paper • 2604.06628 • Published Apr 8 • 330
GrandCode: Achieving Grandmaster Level in Competitive Programming via Agentic Reinforcement Learning Paper • 2604.02721 • Published Apr 3 • 639