Research · Language model reasoning

Language model reasoning

Training language models to reason more effectively, efficiently, and reliably.

Overview

Language-model reasoning is a sequential decision problem in miniature: a model chooses one token at a time, receives a signal about the completed solution, and must balance exploration, accuracy, length, and compute. The lab’s recent work studies how to make that loop more sample-efficient and more reliable, especially when training and inference engines disagree or when the reward arrives only after a full reasoning trace.

Themes

  • Reasoning distillation — transferring useful problem-solving behavior from a stronger teacher into a smaller student without wasting rollout or training compute
  • Rollout efficiency — reusing trajectories, allocating sampling budgets, and selecting informative examples so each generated trace contributes more learning signal
  • Training-inference mismatch — correcting the probability differences that arise when one engine generates a rollout and another engine computes the update
  • Reliable termination — treating equivalent end-of-sequence signals consistently so models do not inflate response length or truncate otherwise-correct solutions
  • Monitorable and compositional reasoning — shaping reasoning traces so they can be inspected, supervised, and recombined across tasks

Open questions

  • How to measure the reasoning value of a trajectory before spending the cost of training on it
  • Which forms of off-policy reuse preserve useful exploration rather than narrowing policy diversity
  • How to separate model capability from artifacts of the training and inference stack
  • How to evaluate reasoning quality, efficiency, and reliability together rather than optimizing one at the expense of the others
  • LLMs
  • Reasoning
  • RL

Publications

Related papers

2026

REVO: Rollout-Efficient Off-Policy Distillation via Variance-Guided Reuse

Yuxiao Yang, Shangzhe Li, Tianrun Yu, Kaixiang Zhao, Taylor W. Killian, Weitong Zhang

arXiv pre-print

REVO makes on-policy distillation more rollout-efficient by reusing student trajectories for multiple learner updates, correcting the resulting policy mismatch with stabilized prefix weighting, one-step resampling, and variance-guided token selection.

An RL View of OPD: Least Square Policy Distillation for Sample-Efficient LLM Reasoning

Shangzhe Li, Yuxiao Yang, Tianrun Yu, Kaixiang Zhao, Xiaoyun Wang, Taylor W. Killian, Weitong Zhang

arXiv pre-print

LSPD reframes on-policy distillation through reinforcement learning, combining optimistic exploration and off-policy trajectory reuse to preserve policy diversity while improving rollout efficiency and mathematical reasoning performance.

Rethinking Training-Inference Mismatch in LLM Reinforcement Learning: Where It Arises and How to Correct It

Tianrun Yu, Kaixiang Zhao, Shangzhe Li, Yuxiao Yang, Porter Jenkins, Weitong Zhang, Taylor W. Killian

arXiv pre-print

CIS corrects the probability mismatch between inference and training engines in RL with verifiable rewards by separating token confidence from engine discrepancy and applying confidence-aware importance-ratio truncation.

When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation

Yuxiao Yang, Tianrun Yu, Shangzhe Li, Kaixiang Zhao, Xuchao Zhang, Chetan Bansal, Huaxiu Yao, Taylor W. Killian, Weitong Zhang

arXiv pre-print

On-policy distillation can suppress a student's preferred end-of-sequence token when the teacher uses a different but equivalent stopping token, causing severe length inflation. Aggregating equivalent EOS probabilities into one semantic stopping action substantially reduces excessive responses and truncation across Qwen3, Llama, and Gemma.

IsoCompute Playbook: Optimally Scaling Sampling Compute for LLM RL

Zhoujun Cheng, Yutao Xie, Yuxiao Qu, Amrith Setlur, Shibo Hao, Varad Pimpalkhute, Tongtong Liang, Feng Yao, Zhengzhong Liu, Eric Xing, Virginia Smith, Ruslan Salakhutdinov, Zhiting Hu, Taylor W. Killian, Aviral Kumar

ICML 2026

Scaling laws for LLM reinforcement learning: how to allocate a fixed compute budget across parallel rollouts, problem batch size, and update steps. Increasing rollouts per problem turns out to be the primary driver, improving solution quality on easy tasks and coverage on hard ones.

All publications