AgentGarten Builds Interactive Worlds for Agents
- •AgentGarten combines simulators, game engines and a shared neural renderer for interactive agent training.
- •Its agents learned hide-and-seek tactics by round 4 and round 10, depending on their role.
- •Researchers report learning from 4 rounds versus millions for a conventional reinforcement learning counterpart.
Researchers introduced AgentGarten, a framework for training agents in real-time interactive virtual worlds, in a paper published on October 8 and submitted to Hugging Face Papers on October 9. The system combines simulators and game engines, which maintain each world’s state and rules, with a shared neural renderer that creates the visual observations agents use. The authors say training environments need consistent state, rules and dynamics as well as observations resembling real-world visual distributions.
AgentGarten adapts a pretrained video model to geometry conditions, then distills it using Adversarial Forcing, a method that makes history prefilling differentiable through exact replay. This lets errors in later predictions update how the renderer encodes earlier observations; real-data adversarial supervision is also used to improve visual quality. The renderer receives structured conditions through a common interface, while simulation backends execute program-defined interaction rules.
Agents perceive rendered views, act in the environment in real time and turn each round of experience into playbooks that later agents inherit and refine. In a hide-and-seek example, hiders build shelters by round 4 and seekers use ramps by round 10. The authors report that agents learned from 4 rounds, compared with millions of rounds for a conventional reinforcement learning counterpart. Because new worlds can be written as code and rendered through the same interface, the framework allows environments to increase in number and difficulty alongside agents.