Anthropic reported extracting millions of human-interpretable features from the middle layer of Claude 3 Sonnet, calling it the first detailed look inside a modern, production-grade large language model. The features ranged from concrete things such as San Francisco and Rosalind Franklin to abstractions such as bugs in code, gender bias in professions, and keeping secrets, with some responding across images and multiple languages. Artificially amplifying a feature changed the model's behavior: turning up the Golden Gate Bridge feature made Claude describe itself as the bridge, and amplifying a scam-email feature overrode its safety training. The work pointed both to a route for monitoring dangerous capabilities and to a means of subverting safeguards. A public demo of the steered model, "Golden Gate Claude," brought interpretability research to a wide audience.