Agents Evaluated Beyond Final Scores
- •Seven frontier models evaluated across 36 long-horizon AI research and engineering tasks
- •Framework measures Solution Framing, Execution, Feedback Control and experience reuse beyond final scores
- •Agents acted more like engineering optimizers than fully autonomous researchers, with rare methodological novelty
Yiwei Li, Wanli Yang, Hexiang Tan and co-authors published “Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development” on Aug 13, with Hugging Face listing the paper on 2026-08-18 after Wanli Yang submitted it on Aug 17. The paper evaluates seven frontier models on 36 long-horizon tasks to test how autonomous agents handle AI research and engineering work that requires extended experimentation rather than one-step answers.
The authors argue that final scores alone do not show where an agent gains or loses progress, or whether earlier experience improves later decisions. Their framework uses rule-based metrics (fixed scoring rules rather than human judgment) to examine within-run behavior across Solution Framing, Execution and Feedback Control, while controlled comparisons measure experience reuse within and across tasks.
The study finds that current agents behave more like engineering optimizers than fully autonomous researchers. They can frame and implement practical solutions, but performance changes substantially across runs, the strongest solutions mostly adapt or combine established techniques, and genuine methodological novelty remains rare.
The paper reports that similar final outcomes can come from different process bottlenecks, meaning two agents with comparable scores may fail at different stages. Experience reuse can help later decisions, but it can also mislead agents, and harness designs (test setups controlling agent tasks) affect performance stability.
The authors point to model training, inference-time strategies, experience management and harness design as areas needing improvement. The Hugging Face page marked the work as #3 Paper of the day, showed 38 upvotes, and listed 0 models, 0 datasets, 0 Spaces and 0 collections citing or including the paper at the time shown.