Scaling Laws for Neural Language Models paper published
OpenAI researchers including Jared Kaplan and Sam McCandlish published "Scaling Laws for Neural Language Models" on arXiv. The paper showed empirically that language model loss improves as a power law with model size, dataset size, and compute, holding across more than seven orders of magnitude. It also found that larger models are more sample-efficient, giving quantitative grounding to the "bigger is better" intuition. These findings underpinned the scaling strategy behind GPT-3 later that year.