2021

December

1 events

Editors' summary

Anthropic puts the criteria it aims at into a paper

On 1 December, Anthropic published its first major paper, framing the evaluation of a general language assistant around three criteria: helpful, honest and harmless.

Six months after founding, the company set out what it would count as good before saying much about what it would build.

This block is written by the editors. It is kept separate from the sourced record below.

Record1 events
  1. 1
    ResearchAnthropic / Amanda Askell / Jared Kaplan / Dario Amodei

    Anthropic's first major paper sets out the HHH criteria

    Anthropic published its first major paper, "A General Language Assistant as a Laboratory for Alignment," framing a general-purpose language assistant as a testbed for alignment research. It distilled the properties of a desirable assistant into three words — helpful, honest, and harmless — and reported that ranked preference modeling substantially outperformed imitation learning and scaled more favorably with model size. Its distinguishing move was to write down what is wanted from an assistant as three properties rather than as performance.