OpenAI announced a multi-year agreement with Cerebras to build 750 megawatts of the company's wafer-scale systems into its inference stack, deployed in stages from 2026. Cerebras said large language models on its hardware return responses up to fifteen times faster than on GPUs, and called it the largest high-speed inference deployment in the world. The capacity was aimed at workloads where latency shapes the experience — coding agents, voice conversation. OpenAI had been widening its compute suppliers across NVIDIA, AMD and Broadcom, but this was the first deal framed around speed of response rather than scale of training. At the time, with the conversation centred on how much training compute a lab could amass, buying latency as a separate line was a different kind of move.