TRACE Speeds FP4 Reinforcement Learning Rollouts
- •TRACE guides FP4 training decisions with quantization results from the rollout path
- •Evaluation covered four MoE language models on reasoning, coding and long-horizon RL tasks
- •Authors report comparable RL performance to BF16 rollouts and up to 5.4x rollout speedup
Researchers Xin Wang, Hao Yu, Zhengyang Zhuge and co-authors propose TRACE, an FP4 quantization framework for reinforcement learning (RL) training of Mixture-of-Experts language models. Their paper, published on arXiv as 2610.07767 and listed on Hugging Face on October 6, addresses the computation and memory overhead of generating rollouts—model outputs used during RL training. Existing FP4 methods optimize quantization accuracy on training and rollout paths separately, rather than directly reducing differences between the two paths.
TRACE uses rollout-guided quantization-aware training: quantization results from the rollout path guide FP4 rounding decisions on the training path. The framework also caches selected mantissa and scale information from deeper layers to reduce the storage and communication overhead of this guidance. The authors evaluated TRACE on four large-scale MoE language models across reasoning, coding and long-horizon RL tasks.
The researchers report that TRACE supports FP4 weights and activations alongside FP4 KV-cache rollouts, with RL performance comparable to BF16 rollouts. It achieved up to 5.4x rollout speedup and stronger final FP4 performance than post-hoc FP4 quantization applied to policies trained in BF16. The paper was submitted by Daniel Wang to Hugging Face on October 7 and was ranked #3 Paper of the Day there; the page showed 38 upvotes at publication time in the provided listing.