-
DataComp-VLM: Improved Open Datasets for Vision-Language Models
Paper • 2606.28551 • Published • 51 -
SkillCoach: Self-Evolving Rubrics for Evaluating and Enhancing Agentic Skill-Use
Paper • 2607.01874 • Published • 21 -
PACE: A Proxy for Agentic Capability Evaluation
Paper • 2607.02032 • Published • 18 -
Measuring the Gap Between Human and LLM Research Ideas
Paper • 2607.01233 • Published • 19
Qiyuan Zhang
DonJoey
AI & ML interests
None yet
Organizations
None yet
Mix-GRM
We provide a collection about ``Beyond Length Scaling: Synergizing Breadth and Depth for
Generative Reward Models'', including data, models, and paper
reading-buffer
-
DataComp-VLM: Improved Open Datasets for Vision-Language Models
Paper • 2606.28551 • Published • 51 -
SkillCoach: Self-Evolving Rubrics for Evaluating and Enhancing Agentic Skill-Use
Paper • 2607.01874 • Published • 21 -
PACE: A Proxy for Agentic Capability Evaluation
Paper • 2607.02032 • Published • 18 -
Measuring the Gap Between Human and LLM Research Ideas
Paper • 2607.01233 • Published • 19
RubricBench
Mix-GRM
We provide a collection about ``Beyond Length Scaling: Synergizing Breadth and Depth for
Generative Reward Models'', including data, models, and paper