DeepMind published "Mastering the game of Go without human knowledge" in Nature, introducing AlphaGo Zero. Given only the rules and no human game records, it learned entirely by reinforcement learning from self-play, starting from random moves and surpassing human level within days across five million self-play games — ending up stronger than the version that beat Lee Sedol. The finding that withholding human game records made the system stronger, not weaker, was received with surprise.