Anthropic's first major paper sets out the HHH criteria
Anthropic published its first major paper, "A General Language Assistant as a Laboratory for Alignment," framing a general-purpose language assistant as a testbed for alignment research. It distilled the properties of a desirable assistant into three words — helpful, honest, and harmless — and reported that ranked preference modeling substantially outperformed imitation learning and scaled more favorably with model size. Its distinguishing move was to write down what is wanted from an assistant as three properties rather than as performance.