John Schulman

Research Scientist at OpenAI

Featured Co-authors

Research Publications

Let's Verify Step by Step 05/31/2023

In recent years, large language models have greatly improved in their ab...

Scaling laws for single-agent reinforcement learning 01/31/2023

Recent work has shown that, in generative modeling, cross-entropy loss i...

Scaling Laws for Reward Model Overoptimization 10/19/2022

In reinforcement learning from human feedback, it is common to optimize...

Efficient Training of Language Models to Fill in the Middle 07/28/2022

We show that autoregressive language models can learn to infill text aft...

Training language models to follow instructions with human feedback 03/04/2022

Making language models bigger does not inherently make them better at...

WebGPT: Browser-assisted question-answering with human feedback 12/17/2021

We fine-tune GPT-3 to answer long-form questions using a text-based web-...

Training Verifiers to Solve Math Word Problems 10/27/2021

State-of-the-art language models can match human performance on many tas...

Batch size-invariance for policy optimization 10/01/2021

We say an algorithm is batch size-invariant if changes to the batch size...

Unsolved Problems in ML Safety 09/28/2021

Machine learning (ML) systems are rapidly increasing in size, are acquir...

Measuring Sample Efficiency and Generalization in Reinforcement Learning Benchmarks: NeurIPS 2020 Procgen Benchmark 03/29/2021

The NeurIPS 2020 Procgen Competition was designed as a centralized bench...

The MineRL 2020 Competition on Sample Efficient Reinforcement Learning using Human Priors 01/26/2021

Although deep reinforcement learning has led to breakthroughs in many di...

Scaling Laws for Autoregressive Generative Modeling 10/28/2020

We identify empirical scaling laws for the cross-entropy loss in four do...

Phasic Policy Gradient 09/09/2020

We introduce Phasic Policy Gradient (PPG), a reinforcement learning fram...

Leveraging Procedural Generation to Benchmark Reinforcement Learning 12/03/2019

In this report, we introduce Procgen Benchmark, a suite of 16 procedural...

Policy Gradient Search: Online Planning and Expert Iteration without Search Trees 04/07/2019

Monte Carlo Tree Search (MCTS) algorithms perform simulation-based searc...

Semi-Supervised Learning by Label Gradient Alignment 02/06/2019

We present label gradient alignment, a novel algorithm for semi-supervis...

Quantifying Generalization in Reinforcement Learning 12/06/2018

In this paper, we investigate the problem of overfitting in deep reinfor...

Model-Based Reinforcement Learning via Meta-Policy Optimization 09/14/2018

Model-based reinforcement learning approaches carry the promise of being...

Gotta Learn Fast: A New Benchmark for Generalization in RL 04/10/2018

In this report, we present a new reinforcement learning (RL) benchmark b...

Reptile: a Scalable Metalearning Algorithm 03/08/2018

This paper considers metalearning problems, where there is a distributio...

Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations 09/28/2017

Dexterous multi-fingered hands are extremely versatile and provide a gen...

Teacher-Student Curriculum Learning 07/01/2017

We propose Teacher-Student Curriculum Learning (TSCL), a framework for a...

UCB Exploration via Q-Ensembles 06/05/2017

We show how an ensemble of Q*-functions can be leveraged for more effec...

#Exploration: A Study of Count-Based Exploration for Deep Reinforcement Learning 11/15/2016

Count-based exploration algorithms are known to perform near-optimally w...

RL²: Fast Reinforcement Learning via Slow Reinforcement Learning 11/09/2016

Deep reinforcement learning (deep RL) has been successful in learning so...

Variational Lossy Autoencoder 11/08/2016

Representation learning seeks to expose certain aspects of observed data...

Concrete Problems in AI Safety 06/21/2016

Rapid progress in machine learning and artificial intelligence (AI) has...

InfoGAN: Interpretable Representation Learning by Information Maximizing Generative Adversarial Nets 06/12/2016

This paper describes InfoGAN, an information-theoretic extension to the...

OpenAI Gym 06/05/2016

OpenAI Gym is a toolkit for reinforcement learning research. It includes...

VIME: Variational Information Maximizing Exploration 05/31/2016

Scalable and effective exploration remains a key challenge in reinforcem...

Theano: A Python framework for fast computation of mathematical expressions 05/09/2016

Theano is a Python library that allows to define, optimize, and evaluate...