OpenAI, in collaboration with DeepMind's safety team, announced the paper "Deep reinforcement learning from human preferences." The method lets an agent infer and learn a goal simply from a human choosing which of two behaviors is better, and was demonstrated by teaching a simulated agent to backflip with about 900 pieces of feedback. An early paper on reinforcement learning from human feedback (RLHF), it offered human preference as a substitute for reward functions that are hard to write down.