Anthropic released Claude 3.5 Sonnet, a mid-tier model that outperformed its own top-end Claude 3 Opus while keeping mid-tier pricing and speed. The company claimed leading scores on graduate-level reasoning (GPQA), undergraduate knowledge (MMLU), and coding (HumanEval), and said internal evaluation had it solving 64% of coding problems against 38% for Claude 3 Opus. Alongside it came Artifacts on claude.ai, which showed generated code and documents in a dedicated pane beside the conversation for viewing and editing — a shift from talking to an AI toward working with one, and an idea widely copied in other chat interfaces. Pricing was $3 per million input tokens and $15 per million output tokens, with a 200K context window.