Policy learning and evaluation from fixed logs of prior decisions, without further interaction.
Learning to act from data someone else collected, for reasons you cannot fully reconstruct, in a world you cannot go back and query.
Identifying states from which an acceptable outcome is no longer reachable, and detecting them early.
States from which good outcomes are no longer reachable — how to recognize them, and how early you can be warned.
Credit assignment when the outcome arrives late, arrives noisy, or never arrives at all.
Credit assignment when outcomes arrive long after the decision that caused them — or arrive too corrupted to trust.
- RL
- Delayed Feedback
- Scientific Discovery
Combinatorial and factored action spaces, where the joint action set is far too large to enumerate.
Decisions composed of many interacting parts, where the number of joint actions is astronomically larger than anything you can enumerate.
The applied settings that generate the lab's methodological problems — clinical care and automated experimentation.
The domains that generate the rest of our problems — where context-sensitive choices carry direct consequences.
- Healthcare
- Clinical
- Scientific Discovery
Training language models to reason more effectively, efficiently, and reliably.
Reinforcement learning, distillation, and inference-aware training methods that improve how language models solve multi-step problems.
Choosing the training examples that make a model learn the right things.
Data selection and trajectory curation for efficient mid-training, where coverage and learnability matter as much as example quality.
- LLMs
- Data Curation
- Mid-training