OpenAI published documentation for GPTBot, the web crawler it uses to gather data for training future models. Pages behind paywalls, pages that gather personally identifiable information, and pages containing text that violates its policies would be filtered out, the company said. Site owners could refuse the crawler by naming GPTBot in robots.txt or by blocking its IP addresses. Publishers and creators began adding the block within days of the documentation appearing.