Compare AIFind AIAI NewsAI How-To
About Us
PrivacyTermsFAQContactContact
AIB Inc.Company info
© 2026 AIB Inc.

Learn2Play Bench Tests Agent Learning

Learn2Play Bench Tests Agent Learning

HuggingFace·Friday, October 9, 2026
  • •National University of Singapore researchers introduce Learn2Play Bench, a set of unfamiliar text-based games for testing agent learning.
  • •Complete records of actions and feedback supported more effective learning than summaries of experience as rules or strategies.
  • •Human players reached higher peak scores; changing an agent harness improved performance and reduced estimated inference cost.
  • •National University of Singapore researchers introduce Learn2Play Bench, a set of unfamiliar text-based games for testing agent learning.
  • •Complete records of actions and feedback supported more effective learning than summaries of experience as rules or strategies.
  • •Human players reached higher peak scores; changing an agent harness improved performance and reduced estimated inference cost.
  • •National University of Singapore researchers introduce Learn2Play Bench, a set of unfamiliar text-based games for testing agent learning.
  • •Complete records of actions and feedback supported more effective learning than summaries of experience as rules or strategies.
  • •Human players reached higher peak scores; changing an agent harness improved performance and reduced estimated inference cost.
  • •National University of Singapore researchers introduce Learn2Play Bench, a set of unfamiliar text-based games for testing agent learning.
  • •Complete records of actions and feedback supported more effective learning than summaries of experience as rules or strategies.
  • •Human players reached higher peak scores; changing an agent harness improved performance and reduced estimated inference cost.

Researchers from the National University of Singapore introduced Learn2Play Bench, a benchmark of text-based games designed to test whether large language model agents can learn from experience in unfamiliar environments. The games use novel or counterintuitive rules, so agents must acquire knowledge by interacting with them rather than relying only on what they learned before. The benchmark provides reproducible feedback and automatic scoring across repeated attempts, and varies game instances to test whether agents apply what they learned in new situations.

The researchers report three findings. Keeping complete records of actions and feedback supported more effective learning than summarizing experience as rules or strategies. Top-performing human players reached higher peak scores than the evaluated agents; humans also explored more varied strategies and repeated actions less often. With the underlying model held fixed, changing the agent harness (the software setup that runs an agent) improved performance while reducing estimated inference cost.

The paper evaluates how underlying models, methods that let agents improve over time, and agent harnesses affect learning ability. Its authors are Yibo Li, Jinhang Qiu, Zhi Zheng, Qianyun Guo, Jiaying Wu, Shuo Ji and Bryan Hooi. The paper was published on October 8, and the Hugging Face Papers page lists it as submitted by Zhi Zheng on October 9 and as its number-one paper of the day.

Researchers from the National University of Singapore introduced Learn2Play Bench, a benchmark of text-based games designed to test whether large language model agents can learn from experience in unfamiliar environments. The games use novel or counterintuitive rules, so agents must acquire knowledge by interacting with them rather than relying only on what they learned before. The benchmark provides reproducible feedback and automatic scoring across repeated attempts, and varies game instances to test whether agents apply what they learned in new situations.

The researchers report three findings. Keeping complete records of actions and feedback supported more effective learning than summarizing experience as rules or strategies. Top-performing human players reached higher peak scores than the evaluated agents; humans also explored more varied strategies and repeated actions less often. With the underlying model held fixed, changing the agent harness (the software setup that runs an agent) improved performance while reducing estimated inference cost.

The paper evaluates how underlying models, methods that let agents improve over time, and agent harnesses affect learning ability. Its authors are Yibo Li, Jinhang Qiu, Zhi Zheng, Qianyun Guo, Jiaying Wu, Shuo Ji and Bryan Hooi. The paper was published on October 8, and the Hugging Face Papers page lists it as submitted by Zhi Zheng on October 9 and as its number-one paper of the day.

Read original (English)·Oct 9, 2026
#learn2play bench#llm agents#agent learning#text based games#experience retention#agent harness#inference cost#national university of singapore