InstructGPT unveiled; RLHF-tuned models become API default
OpenAI unveiled InstructGPT, versions of GPT-3 fine-tuned to follow instructions using reinforcement learning from human feedback (RLHF), and deployed them as the default language models on its API. Labelers preferred outputs from the 1.3-billion-parameter InstructGPT over those of the 175-billion-parameter GPT-3, a model more than 100x larger. It was the first large-scale production deployment of alignment research and laid the technical groundwork for ChatGPT later that year.