Anthropic released an upgraded Claude 3.5 Sonnet and a new Claude 3.5 Haiku, and put "computer use" into public beta. Instead of purpose-built tools, Claude could now look at a screen, move a cursor, click buttons, and type — operating software the way a person does. Anthropic conceded it was experimental, "at times cumbersome and error-prone," and said it had shipped early to get developer feedback. On OSWorld, Claude scored 14.9% in the screenshot-only condition against 7.8% for the next-best system, rising to 22.0% with additional interaction steps. The upgraded Sonnet improved on SWE-bench Verified from 33.4% to 49.0%, which Anthropic said led all publicly available models including OpenAI's o1-preview.