# John Schulman

Research Scientist at OpenAI

## Featured Co-authors

- [Yoshua Bengio](/content/profile/yoshua-bengio/index.html) 448 publications
- [Sergey Levine](/content/profile/sergey-levine/index.html) 379 publications
- [Xi Chen](/content/profile/xi-chen/index.html) 293 publications
- [Pieter Abbeel](/content/profile/pieter-abbeel/index.html) 263 publications
- [Kyunghyun Cho](/content/profile/kyunghyun-cho/index.html) 218 publications
- [Ruslan Salakhutdinov](/content/profile/ruslan-salakhutdinov/index.html) 194 publications
- [Aaron Courville](/content/profile/aaron-courville/index.html) 152 publications
- [Jie Tang](/content/profile/jie-tang/index.html) 135 publications
- [Abhishek Gupta](/content/profile/abhishek-gupta/index.html) 133 publications
- [Ying Zhang](/content/profile/ying-zhang/index.html) 107 publications
- [Oriol Vinyals](/content/profile/oriol-vinyals/index.html) 101 publications

## Research Publications

### [Let's Verify Step by Step](/content/publication/let-s-verify-step-by-step/index.html) 05/31/2023
In recent years, large language models have greatly improved in their ab...

### [Scaling laws for single-agent reinforcement learning](/content/publication/scaling-laws-for-single-agent-reinforcement-learning/index.html) 01/31/2023
Recent work has shown that, in generative modeling, cross-entropy loss i...

### [Scaling Laws for Reward Model Overoptimization](/content/publication/scaling-laws-for-reward-model-overoptimization/index.html) 10/19/2022
In reinforcement learning from human feedback, it is common to optimize...

### [Efficient Training of Language Models to Fill in the Middle](/content/publication/efficient-training-of-language-models-to-fill-in-the-middle/index.html) 07/28/2022
We show that autoregressive language models can learn to infill text aft...

### [Training language models to follow instructions with human feedback](/content/publication/training-language-models-to-follow-instructions-with-human-feedback/index.html) 03/04/2022
Making language models bigger does not inherently make them better at...

### [WebGPT: Browser-assisted question-answering with human feedback](/content/publication/webgpt-browser-assisted-question-answering-with-human-feedback/index.html) 12/17/2021
We fine-tune GPT-3 to answer long-form questions using a text-based web-...

### [Training Verifiers to Solve Math Word Problems](/content/publication/training-verifiers-to-solve-math-word-problems/index.html) 10/27/2021
State-of-the-art language models can match human performance on many tas...

### [Batch size-invariance for policy optimization](/content/publication/batch-size-invariance-for-policy-optimization/index.html) 10/01/2021
We say an algorithm is batch size-invariant if changes to the batch size...

### [Unsolved Problems in ML Safety](/content/publication/unsolved-problems-in-ml-safety/index.html) 09/28/2021
Machine learning (ML) systems are rapidly increasing in size, are acquir...

### [Measuring Sample Efficiency and Generalization in Reinforcement Learning Benchmarks: NeurIPS 2020 Procgen Benchmark](/content/publication/measuring-sample-efficiency-and-generalization-in-reinforcement-learning-benchmarks-neurips-2020-procgen-benchmark/index.html) 03/29/2021
The NeurIPS 2020 Procgen Competition was designed as a centralized bench...

### [The MineRL 2020 Competition on Sample Efficient Reinforcement Learning using Human Priors](/content/publication/the-minerl-2020-competition-on-sample-efficient-reinforcement-learning-using-human-priors/index.html) 01/26/2021
Although deep reinforcement learning has led to breakthroughs in many di...

### [Scaling Laws for Autoregressive Generative Modeling](/content/publication/scaling-laws-for-autoregressive-generative-modeling/index.html) 10/28/2020
We identify empirical scaling laws for the cross-entropy loss in four do...

### [Phasic Policy Gradient](/content/publication/phasic-policy-gradient/index.html) 09/09/2020
We introduce Phasic Policy Gradient (PPG), a reinforcement learning fram...

### [Leveraging Procedural Generation to Benchmark Reinforcement Learning](/content/publication/leveraging-procedural-generation-to-benchmark-reinforcement-learning/index.html) 12/03/2019
In this report, we introduce Procgen Benchmark, a suite of 16 procedural...

### [Policy Gradient Search: Online Planning and Expert Iteration without Search Trees](/content/publication/policy-gradient-search-online-planning-and-expert-iteration-without-search-trees/index.html) 04/07/2019
Monte Carlo Tree Search (MCTS) algorithms perform simulation-based searc...

### [Semi-Supervised Learning by Label Gradient Alignment](/content/publication/semi-supervised-learning-by-label-gradient-alignment/index.html) 02/06/2019
We present label gradient alignment, a novel algorithm for semi-supervis...

### [Quantifying Generalization in Reinforcement Learning](/content/publication/quantifying-generalization-in-reinforcement-learning/index.html) 12/06/2018
In this paper, we investigate the problem of overfitting in deep reinfor...

### [Model-Based Reinforcement Learning via Meta-Policy Optimization](/content/publication/model-based-reinforcement-learning-via-meta-policy-optimization/index.html) 09/14/2018
Model-based reinforcement learning approaches carry the promise of being...

### [Gotta Learn Fast: A New Benchmark for Generalization in RL](/content/publication/gotta-learn-fast-a-new-benchmark-for-generalization-in-rl/index.html) 04/10/2018
In this report, we present a new reinforcement learning (RL) benchmark b...

### [Reptile: a Scalable Metalearning Algorithm](/content/publication/reptile-a-scalable-metalearning-algorithm/index.html) 03/08/2018
This paper considers metalearning problems, where there is a distributio...

### [Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations](/content/publication/learning-complex-dexterous-manipulation-with-deep-reinforcement-learning-and-demonstrations/index.html) 09/28/2017
Dexterous multi-fingered hands are extremely versatile and provide a gen...

### [Teacher-Student Curriculum Learning](/content/publication/teacher-student-curriculum-learning/index.html) 07/01/2017
We propose Teacher-Student Curriculum Learning (TSCL), a framework for a...

### [UCB Exploration via Q-Ensembles](/content/publication/ucb-exploration-via-q-ensembles/index.html) 06/05/2017
We show how an ensemble of Q*-functions can be leveraged for more effec...

### [#Exploration: A Study of Count-Based Exploration for Deep Reinforcement Learning](/content/publication/exploration-a-study-of-count-based-exploration-for-deep-reinforcement-learning/index.html) 11/15/2016
Count-based exploration algorithms are known to perform near-optimally w...

### [RL²: Fast Reinforcement Learning via Slow Reinforcement Learning](/content/publication/rl-2-fast-reinforcement-learning-via-slow-reinforcement-learning/index.html) 11/09/2016
Deep reinforcement learning (deep RL) has been successful in learning so...

### [Variational Lossy Autoencoder](/content/publication/variational-lossy-autoencoder/index.html) 11/08/2016
Representation learning seeks to expose certain aspects of observed data...

### [Concrete Problems in AI Safety](/content/publication/concrete-problems-in-ai-safety/index.html) 06/21/2016
Rapid progress in machine learning and artificial intelligence (AI) has...

### [InfoGAN: Interpretable Representation Learning by Information Maximizing Generative Adversarial Nets](/content/publication/infogan-interpretable-representation-learning-by-information-maximizing-generative-adversarial-nets/index.html) 06/12/2016
This paper describes InfoGAN, an information-theoretic extension to the...

### [OpenAI Gym](/content/publication/openai-gym/index.html) 06/05/2016
OpenAI Gym is a toolkit for reinforcement learning research. It includes...

### [VIME: Variational Information Maximizing Exploration](/content/publication/vime-variational-information-maximizing-exploration/index.html) 05/31/2016
Scalable and effective exploration remains a key challenge in reinforcem...

### [Theano: A Python framework for fast computation of mathematical expressions](/content/publication/theano-a-python-framework-for-fast-computation-of-mathematical-expressions/index.html) 05/09/2016
Theano is a Python library that allows to define, optimize, and evaluate...
