Image GPT
Image GPT: transformer trained on pixel sequences for image completion and generation with competitive unsupervised classification.
Image GPT: transformer trained on pixel sequences for image completion and generation with competitive unsupervised classification.
Procgen Benchmark: 16 procedurally-generated environments to measure reinforcement learning agent generalization and learning speed.
Safety Gym: open-source suite of environments and tools for measuring reinforcement learning agent progress with safety constraints.
Final GPT-2 1.5B model release with code/weights and staged release process documentation as test case for future powerful models.
Fine-tuned 774M GPT-2 using human feedback for summarization and text continuation tasks, achieving preference alignment with 60k human labels.
Multi-agent reinforcement learning agents discover six emergent strategies in hide-and-seek game, demonstrating self-supervised emergent complexity.
Release of 774M parameter GPT-2 language model with staged release strategy, open-source legal agreements, and research on misuse/benefit.
Policy research paper on four strategies for industry cooperation on AI safety: risk communication, technical collaboration, transparency, standards.
OpenAI Fellows program conclusion with 6-month apprenticeship projects. Career/education program announcement.
Massively multiagent game environment for RL supporting variable agent counts with persistent open-ended tasks.
Paper arguing AI safety alignment research needs social scientists to address human psychology and value alignment.
Large-scale unsupervised language model achieving SOTA on multiple NLP tasks without task-specific training.
OpenAI Fellows apprenticeship program completion announcement with participant transitions to core contributor roles.
Analysis showing gradient noise scale predicts neural network training parallelizability across diverse tasks.
CoinRun environment benchmarks agent generalization in RL; gradient noise scale predicts training parallelizability.
Educational resource with code examples, exercises, and tutorials for learning deep reinforcement learning.
Energy-based model learns spatial concepts from few demonstrations with cross-domain transfer to robotics tasks.
Model-based control approach for efficient learning and exploration combining online planning with offline learning.
Random Network Distillation prediction-based method for curiosity-driven exploration, first to exceed human performance on Montezuma's Revenge.
Iterated amplification technique for AI safety enabling specification of complex goals by decomposing tasks into simpler sub-tasks.
Large-scale study examining curiosity-driven learning approaches in machine learning agents.
OpenAI Five defeats 99.95th percentile Dota 2 players in best-of-three match with live audience and 100k viewers.
Robot hand trained to manipulate physical objects with unprecedented dexterity using machine learning methods.
OpenAI Five benchmark match announcement with removed gameplay restrictions for competition against professional Dota players.
Agent trained to score 74,500 on Montezuma's Revenge from single demonstration using PPO from carefully chosen state sequences.
Results from Retro Contest showing top performers used tuning/extensions of PPO and Rainbow algorithms on Sonic benchmark.
General learning framework for modeling agent behavior in multiagent systems using representation learning from minimal interaction data.
State-of-the-art language understanding results using transformers with unsupervised pre-training, with code release for reproducibility.
GamePad system for exploring machine learning methods applied to theorem proving in Coq proof assistant using step-by-step human supervision.
OpenAI Fellows program accepting applications for 6-month AI research apprenticeship. Targets individuals without formal background in AI with mentorship and team placement.
Analysis showing AI compute in largest training runs increased 300,000x since 2012 with 3.4-month doubling time versus 2-year Moore's Law period.
AI safety technique training agents to debate topics with human judge. Proposes approach to align advanced AI systems with human preferences with proof-of-concept experiments.
RL benchmark based on Sonic the Hedgehog measuring transfer and few-shot learning. Presents baseline algorithms and evaluation methodology.
Transfer learning RL competition measuring generalization from previous experience on unseen video game levels. Uses Gym Retro platform with new benchmark.
Action-dependent factorized baselines for variance reduction in policy gradient methods. Bias-free approach exploiting stochastic policy structure without MDP assumptions.
Analysis of first-order meta-learning algorithms for learning parameter initializations. Generalizes first-order MAML using only first-order derivatives for meta-learning updates.
Reptile: scalable meta-learning algorithm using repeated task sampling and SGD updates. Mathematically similar to first-order MAML requiring only black-box optimizer access.
Two new meta-reinforcement learning algorithms (E-MAML, E-RL²) for exploration problems. Evaluated on Krazy World and maze environments with improved performance.
Technical report introducing multi-goal RL benchmark with continuous control tasks for Fetch robotic arm and Shadow hand. Sparse reward framework integrated with OpenAI Gym.
Method for training AIs to teach each other using human-interpretable examples. Automatically selects informative examples to teach concepts effectively to both AI and human learners.