Google announced Gemini 1.5 Pro with an experimental one-million-token context window, offered in private preview through Google AI Studio, against a 128,000-token default. The largest window on any publicly available model until then had been 200,000 tokens, in Anthropic's Claude 2.1 — an order-of-magnitude jump. Entire books, whole codebases, or hours of video could be read in a single pass, and long context became a clear point of differentiation for Google. The model used a Mixture-of-Experts (MoE) architecture, running only a subset of its internal "expert" networks for any given input, which Google said made it more efficient to train and serve — the first time Google described a flagship model of its own as MoE. Arriving just two and a half months after Gemini 1.0, it showed how fast Google was moving to close the gap. Context length settled in as a standard axis of competition, with Anthropic reaching a million tokens in 2026.