OpenAI Allows Internet Users to Block its Web Crawler, GPTBot, for Site Scraping
Matthew David
News
OpenAI now allows website owners to block its web crawler, GPTBot, from scraping specific content. While scraped data improves models, sensitive or policy-breaking content will be excluded. Allowing GPTBot access enhances AI accuracy, and this step might lead to data opt-out options. The sourcing of training data has led to debates, with platforms like Reddit and Twitter limiting usage, and OpenAI considering content watermarking but not fully committing to ending internet data use.