2017

8 events

Editors' summary

The Transformer paper appears, and AlphaGo leaves Go behind

On 12 June, researchers at Google published the Transformer paper, an architecture built from attention alone, without recurrence or convolution. It improved translation while making training parallel.

In May, AlphaGo beat Ke Jie, the world's top player, three games to none, and retired from Go. In October, AlphaGo Zero surpassed it using no human games at all, learning from self-play. Learning without human play as a model was shown to work.

OpenAI pushed on with reinforcement learning. On 13 June it published work, with Google, on training from nothing more than a human saying which of two behaviours they preferred. July brought PPO, an algorithm easy enough to handle that it came into wide use. In August its bot beat a professional at one-on-one Dota 2. In April it had reported finding, inside a model trained on text without supervision, a unit that tracked sentiment on its own.

On 20 June, Andrej Karpathy left OpenAI to lead AI at Tesla.

This block is written by the editors. It is kept separate from the sourced record below.

Record8 events
  1. Apr 6
    ResearchOpenAI / Alec Radford / Rafał Józefowicz / Ilya Sutskever

    Unsupervised sentiment neuron research announced

    Alec Radford and colleagues at OpenAI announced research showing that a model trained only to predict the next character in Amazon reviews learned an excellent unsupervised representation of sentiment. The model contained a single "sentiment neuron" carrying almost all of the sentiment signal, and achieved a then state-of-the-art 91.8% accuracy on the Stanford Sentiment Treebank. The result bolstered the view that unsupervised next-token prediction could serve as pretraining.

  2. May 27
    ResearchGoogle / Demis Hassabis / Ke Jie / David Silver

    AlphaGo sweeps world No. 1 Ke Jie, then retires

    At the Future of Go Summit in Wuzhen, China, held from 23 to 27 May, DeepMind's AlphaGo Master beat world No. 1 Ke Jie in all three games of their match. The first was decided by half a point, and Ke Jie was moved to tears afterward. DeepMind's Demis Hassabis announced that the summit would be AlphaGo's last competitive event, saying the team would turn to general algorithms aimed at scientific problems — finding cures for diseases, cutting energy consumption, inventing new materials. The question of whether AI could beat the best human Go players was settled.

  3. Jun 12
    ResearchGoogle / Ashish Vaswani / Noam Shazeer / Niki Parmar

    "Attention Is All You Need" introduces the Transformer

    Eight Google researchers published "Attention Is All You Need," proposing a new architecture they called the Transformer. Dispensing with recurrence and convolution entirely and handling sequences through attention alone, it beat prior approaches on two machine translation tasks while being far more parallelizable and much faster to train. The authors were Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. Conceived to improve translation, the architecture became the shared foundation of GPT, BERT, and essentially every large language model since — the single most consequential paper in the history of generative AI, and one whose publication by Google made its competitors' rise possible. In the years that followed, nearly all eight authors left Google to found or join AI companies of their own.

  4. Jun 13
    ResearchOpenAI / Google / Paul Christiano / Jan Leike

    Learning from human preferences research announced

    OpenAI, in collaboration with DeepMind's safety team, announced the paper "Deep reinforcement learning from human preferences." The method lets an agent infer and learn a goal simply from a human choosing which of two behaviors is better, and was demonstrated by teaching a simulated agent to backflip with about 900 pieces of feedback. An early paper on reinforcement learning from human feedback (RLHF), it offered human preference as a substitute for reward functions that are hard to write down.

  5. Jun 20
    PeopleOpenAI / Andrej Karpathy / Elon Musk / Tesla

    Andrej Karpathy leaves OpenAI to lead AI at Tesla

    Andrej Karpathy, a founding member and researcher at OpenAI, joined Tesla as Director of AI, leading AI and Autopilot Vision. Karpathy announced the move on X, and Tesla confirmed it in a statement. The hire, in which Elon Musk recruited from the very organization he backed, drew attention as emblematic of the relationship between the two organizations.

  6. Jul 20
    ResearchOpenAI / John Schulman / Filip Wolski / Prafulla Dhariwal

    PPO reinforcement learning algorithm released

    John Schulman and colleagues at OpenAI released Proximal Policy Optimization (PPO), publishing both the paper and an implementation. PPO performed comparably to or better than existing methods while being much simpler to implement and tune, and had already become OpenAI's default algorithm at the time of release. It went on to become one of the most widely used reinforcement learning methods and later played a central role in fine-tuning language models with RLHF.

  7. Aug 11
    ResearchOpenAI / Danil "Dendi" Ishutin / Syed "SumaiL" Hassan

    Dota 2 bot beats pro player at The International

    An OpenAI bot trained entirely through self-play defeated professional player Dendi in a 1v1 Dota 2 match on the main stage of The International. Learning from scratch without imitation learning or tree search, the bot had also gone undefeated against several top professionals, including SumaiL, in the preceding week. Reaching this level in a complex game that plays out in real time, rather than a turn-based one like Go or chess, was a substantial result at the time.

  8. Oct 19
    ResearchGoogle / David Silver / Julian Schrittwieser / Demis Hassabis

    AlphaGo Zero learns Go from scratch, with no human data

    DeepMind published "Mastering the game of Go without human knowledge" in Nature, introducing AlphaGo Zero. Given only the rules and no human game records, it learned entirely by reinforcement learning from self-play, starting from random moves and surpassing human level within days across five million self-play games — ending up stronger than the version that beat Lee Sedol. The finding that withholding human game records made the system stronger, not weaker, was received with surprise.