Researchers Propose Adaptive Reward Routing
- •Tencent-affiliated researchers propose adaptive reward routing for joint audio-video diffusion training
- •Method adjusts update locations and coordinates competing rewards during forward-process reinforcement learning
- •Authors report improvements in modality quality, semantic consistency, and audio-video synchronization over strong baselines
Tencent-affiliated researchers published Adaptive Reward Routing on September 29, a method for training joint audio-video diffusion models with multiple rewards. The paper was submitted to Hugging Face by Eddie on October 2 and was listed as the platform’s No. 2 Paper of the Day. The authors say the method improves audio and video quality, consistency between the two modalities, and synchronization, although the article gives no numerical results.
The researchers identify two training challenges: deciding where reward-based updates should act and balancing rewards that can compete. Existing methods often use fixed routing and reward weights, which the authors say do not track changes in how the model functions during training. Adaptive Reward Routing adjusts both update locations and reward coordination during forward-process reinforcement learning, the paper’s training approach for diffusion models.
Its first component, Cross-Modal Influence-Guided Routing, uses bidirectional cross-attention responses as a proxy for changing influence between audio and video. It dynamically reweights token-aware losses and scales gradients across cross-modal layers, without additional model interventions. Its second component, Preference-Preserving Modality-Aware Reweighting, keeps predefined reward weights as preference priors and applies branch-specific reward-gradient interactions as residual corrections after warm-up. The authors say this helps address changing conflicts while preventing stronger rewards from suppressing weaker but essential objectives.
Experiments reportedly showed consistent improvements over strong reinforcement-learning baselines in modality quality, semantic consistency, and audio-video synchronization. Ablations and mechanism analyses also supported the separate, complementary contributions of adaptive update routing and reward coordination. The article does not report scores, dataset names, or detailed experiment results.