John Schulman
Research Scientist at OpenAI
Featured Co-authors
- Yoshua Bengio 448 publications
- Sergey Levine 379 publications
- Xi Chen 293 publications
- Pieter Abbeel 263 publications
- Kyunghyun Cho 218 publications
- Ruslan Salakhutdinov 194 publications
- Aaron Courville 152 publications
- Jie Tang 135 publications
- Abhishek Gupta 133 publications
- Ying Zhang 107 publications
- Oriol Vinyals 101 publications
Research Publications
Let's Verify Step by Step 05/31/2023
In recent years, large language models have greatly improved in their ab...
Scaling laws for single-agent reinforcement learning 01/31/2023
Recent work has shown that, in generative modeling, cross-entropy loss i...
Scaling Laws for Reward Model Overoptimization 10/19/2022
In reinforcement learning from human feedback, it is common to optimize...
Efficient Training of Language Models to Fill in the Middle 07/28/2022
We show that autoregressive language models can learn to infill text aft...
Training language models to follow instructions with human feedback 03/04/2022
Making language models bigger does not inherently make them better at...
WebGPT: Browser-assisted question-answering with human feedback 12/17/2021
We fine-tune GPT-3 to answer long-form questions using a text-based web-...
Training Verifiers to Solve Math Word Problems 10/27/2021
State-of-the-art language models can match human performance on many tas...
Batch size-invariance for policy optimization 10/01/2021
We say an algorithm is batch size-invariant if changes to the batch size...
Unsolved Problems in ML Safety 09/28/2021
Machine learning (ML) systems are rapidly increasing in size, are acquir...
Measuring Sample Efficiency and Generalization in Reinforcement Learning Benchmarks: NeurIPS 2020 Procgen Benchmark 03/29/2021
The NeurIPS 2020 Procgen Competition was designed as a centralized bench...
The MineRL 2020 Competition on Sample Efficient Reinforcement Learning using Human Priors 01/26/2021
Although deep reinforcement learning has led to breakthroughs in many di...
Scaling Laws for Autoregressive Generative Modeling 10/28/2020
We identify empirical scaling laws for the cross-entropy loss in four do...
Phasic Policy Gradient 09/09/2020
We introduce Phasic Policy Gradient (PPG), a reinforcement learning fram...
Leveraging Procedural Generation to Benchmark Reinforcement Learning 12/03/2019
In this report, we introduce Procgen Benchmark, a suite of 16 procedural...
Policy Gradient Search: Online Planning and Expert Iteration without Search Trees 04/07/2019
Monte Carlo Tree Search (MCTS) algorithms perform simulation-based searc...
Semi-Supervised Learning by Label Gradient Alignment 02/06/2019
We present label gradient alignment, a novel algorithm for semi-supervis...
Quantifying Generalization in Reinforcement Learning 12/06/2018
In this paper, we investigate the problem of overfitting in deep reinfor...
Model-Based Reinforcement Learning via Meta-Policy Optimization 09/14/2018
Model-based reinforcement learning approaches carry the promise of being...
Gotta Learn Fast: A New Benchmark for Generalization in RL 04/10/2018
In this report, we present a new reinforcement learning (RL) benchmark b...
Reptile: a Scalable Metalearning Algorithm 03/08/2018
This paper considers metalearning problems, where there is a distributio...
Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations 09/28/2017
Dexterous multi-fingered hands are extremely versatile and provide a gen...
Teacher-Student Curriculum Learning 07/01/2017
We propose Teacher-Student Curriculum Learning (TSCL), a framework for a...
UCB Exploration via Q-Ensembles 06/05/2017
We show how an ensemble of Q*-functions can be leveraged for more effec...
#Exploration: A Study of Count-Based Exploration for Deep Reinforcement Learning 11/15/2016
Count-based exploration algorithms are known to perform near-optimally w...
RL²: Fast Reinforcement Learning via Slow Reinforcement Learning 11/09/2016
Deep reinforcement learning (deep RL) has been successful in learning so...
Variational Lossy Autoencoder 11/08/2016
Representation learning seeks to expose certain aspects of observed data...
Concrete Problems in AI Safety 06/21/2016
Rapid progress in machine learning and artificial intelligence (AI) has...
InfoGAN: Interpretable Representation Learning by Information Maximizing Generative Adversarial Nets 06/12/2016
This paper describes InfoGAN, an information-theoretic extension to the...
OpenAI Gym 06/05/2016
OpenAI Gym is a toolkit for reinforcement learning research. It includes...
VIME: Variational Information Maximizing Exploration 05/31/2016
Scalable and effective exploration remains a key challenge in reinforcem...
Theano: A Python framework for fast computation of mathematical expressions 05/09/2016
Theano is a Python library that allows to define, optimize, and evaluate...