OpenAI released o1-preview and o1-mini, preview versions of OpenAI o1, a reasoning model that spends time "thinking" before answering. Trained with reinforcement learning to refine its chain of thought and using more compute at inference time, it substantially outperformed GPT-4o on hard math, coding, and science problems. What was new at the time was the demonstration that performance could be bought with inference-time compute, not only with training scale.