๐๏ธ On Vacation
Yuling
YerbaPage
AI & ML interests
None yet
Recent Activity
authored a paper 1 day ago
SRPO: Self-Reflective Policy Optimization for Long-Horizon Reasoning authored a paper 1 day ago
EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction posted an update 1 day ago
You don't have to run an agent evaluation to the end.
An agent's final outcome is usually visible from its early behavior, so we just stop the run once it's clear.
EarlyEval cuts 13% to 26% of steps and up to 44% of input tokens, and resolve rates move by only 1 to 2 points.
https://huggingface.co/papers/2609.02783
https://github.com/inphotoo/earlyeval