Most publishers allow AI bots to crawl their sites (newoldweb.com)

🤖 AI Summary
A new analysis of robots.txt files from 5,818 English-language media sites finds that most publishers have not broadly blocked AI crawlers: only 32% of sites block at least one AI bot. Blocking rates vary by ownership—13% of non-profits, 33% of private companies and 51% of publicly traded media firms block any AI bot—suggesting larger or centrally managed publishers are more likely to act. Specific user-agent blocking: OpenAI’s GPTBot is disallowed by 29% of sites, Common Crawl’s CCBot by 27%, Google‑Extended by 24%, and Anthropic user agents by about 21%; Perplexity is blocked by 20%. General-interest outlets are the most aggressive (69% block AI bots), local sites about half, while niche verticals like sports and entertainment block the fewest. The findings matter because robots.txt—an honor-system tool long used to shape crawler traffic—is being repurposed as a strategic lever to control whether site content feeds LLM training. Technically, robots.txt can stop future crawls but cannot reliably prevent malicious/skulking bots and cannot “unlearn” content already encoded in neural weights: data ingested by earlier models may persist across model versions. That creates tricky tradeoffs for publishers between discovery/traffic and protecting original reporting or monetization. The study used automated and manual site classification (including ChatGPT-assisted tagging) and a data-science workflow to extract and analyze blocked user agents, highlighting both a shifting industry posture and the limits of current web‑scale governance tools.
Loading comments...
loading comments...