2026
REVO: Rollout-Efficient Off-Policy Distillation via Variance-Guided Reuse
Yuxiao Yang, Shangzhe Li, Tianrun Yu, Kaixiang Zhao, Taylor W. Killian, Weitong Zhang
arXiv pre-print
REVO makes on-policy distillation more rollout-efficient by reusing student trajectories for multiple learner updates, correcting the resulting policy mismatch with stabilized prefix weighting, one-step resampling, and variance-guided token selection.
An RL View of OPD: Least Square Policy Distillation for Sample-Efficient LLM Reasoning
Shangzhe Li, Yuxiao Yang, Tianrun Yu, Kaixiang Zhao, Xiaoyun Wang, Taylor W. Killian, Weitong Zhang
arXiv pre-print
LSPD reframes on-policy distillation through reinforcement learning, combining optimistic exploration and off-policy trajectory reuse to preserve policy diversity while improving rollout efficiency and mathematical reasoning performance.
Rethinking Training-Inference Mismatch in LLM Reinforcement Learning: Where It Arises and How to Correct It
Tianrun Yu, Kaixiang Zhao, Shangzhe Li, Yuxiao Yang, Porter Jenkins, Weitong Zhang, Taylor W. Killian
arXiv pre-print
CIS corrects the probability mismatch between inference and training engines in RL with verifiable rewards by separating token confidence from engine discrepancy and applying confidence-aware importance-ratio truncation.
When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation
Yuxiao Yang, Tianrun Yu, Shangzhe Li, Kaixiang Zhao, Xuchao Zhang, Chetan Bansal, Huaxiu Yao, Taylor W. Killian, Weitong Zhang
arXiv pre-print
On-policy distillation can suppress a student's preferred end-of-sequence token when the teacher uses a different but equivalent stopping token, causing severe length inflation. Aggregating equivalent EOS probabilities into one semantic stopping action substantially reduces excessive responses and truncation across Qwen3, Llama, and Gemma.
IsoCompute Playbook: Optimally Scaling Sampling Compute for LLM RL
Zhoujun Cheng, Yutao Xie, Yuxiao Qu, Amrith Setlur, Shibo Hao, Varad Pimpalkhute, Tongtong Liang, Feng Yao, Zhengzhong Liu, Eric Xing, Virginia Smith, Ruslan Salakhutdinov, Zhiting Hu, Taylor W. Killian, Aviral Kumar
ICML 2026
Scaling laws for LLM reinforcement learning: how to allocate a fixed compute budget across parallel rollouts, problem batch size, and update steps. Increasing rollouts per problem turns out to be the primary driver, improving solution quality on easy tasks and coverage on hard ones.