ProRL Improves Reinforcement Learning for Recommender Systems
- •Fudan University researchers released ProRL to fix gradient estimation errors in proactive recommendation systems.
- •ProRL uses stepwise reward centering and position-specific advantage estimation to significantly reduce gradient variance and bias.
- •The model outperformed state-of-the-art baselines on MovieLens-1M, Steam, and Amazon-Book datasets, according to the paper.
Researchers from Fudan University introduced ProRL, a reinforcement learning (RL) framework designed to improve Proactive Recommender Systems (PRSs). Published on May 27, 2026, and accepted for ICML 2026, the study addresses flaws in how standard policy gradients function when guiding users toward target items through sequences of recommendations. Conventional methods suffer from a length-dependent bias where models favor longer paths simply because they accumulate more reward, leading to repetitive outputs and low diversity. Additionally, these systems often experience high gradient variance because they weight every action by the total path reward, ignoring the fact that early actions do not influence subsequent rewards.
To resolve these issues, ProRL implements two core mechanisms. Stepwise Reward Centering subtracts expected rewards to neutralize the length-dependent bias, preventing the model from exploiting path length to gain higher scores. Position-Specific Advantage Estimation uses step-dependent baselines to reduce gradient variance, achieving approximately 5% of the variance found in the standard REINFORCE algorithm without needing a learned critic. This approach allows the policy to focus on actual recommendation quality rather than mathematical artifacts.
Experimental results across the MovieLens-1M, Steam, and Amazon-Book datasets show that ProRL consistently outperforms existing sequential, heuristic, supervised, and LLM-based baselines across four metrics. The model demonstrated strong generalization, as it succeeded even when evaluated against reward signals not used during training. Furthermore, cross-evaluator testing against models like GRU4Rec, BERT4Rec, and LightSANs confirmed that the guidance strategies learned by ProRL remain effective outside of the specific training environment. The team released the implementation code on GitHub.