China's DeepSeek released DeepSeek-V3, a sparse mixture-of-experts model with 671 billion parameters in total and 37 billion active per token, trained on 14.8 trillion tokens. The run took 2.788 million H800 GPU-hours, which the company put at roughly $5.6 million against contemporary estimates of $50–100 million to train GPT-4 — an order of magnitude less for performance close to the frontier. The weights were published. It was read as a lab cut off from the best GPUs by export controls squeezing efficiency out of the hardware it could get, which raised its own questions about how well the controls were working. Observers noted at the time that the $5.6 million covered the final training run only, not the research, failed attempts, or the cluster itself.