🤖 AI Summary
robots.txt — the 1994 convention-born text file formalized as RFC 9309 in 2022 — has effectively been retired as a universal traffic cop for the web. Decades of voluntary compliance unraveled under the scale of modern AI crawlers: the web’s old crawl-to-referral economics (Google once averaged ~14:1) broke down when AI bots generated massive fetches with almost no downstream referrals. Cloudflare’s July 2025 decision to block AI crawlers by default symbolized the collapse of voluntary norms; measurements showed industrial-scale extraction (OpenAI-linked crawls hitting ~1,700:1 crawl-to-referral, Anthropic’s ClaudeBot in the tens of thousands to one). The Internet Archive’s 2017 choice to ignore robots.txt and tactics like undisclosed IPs by Perplexity further signaled that polite opt-outs no longer constrain large actors.
Technically and socially, the loss matters because it shifts the governance model from a lightweight, interoperable signal to an arms race of enforcement and fragmentation. Site operators are deploying TLS-based crawler fingerprinting, honeypots and behavioral analysis while regulators (EDPB Opinion 28/2024, Italy’s €15M OpenAI fine) push legal remedies. Proposed successors (ai.txt, TDM ReP, No-AI-Training headers) compete with commercial paywalls, licensing deals and network-level blocks, risking a fractured web where access depends on contracts or technical gatekeeping. For the AI/ML community this means harder, messier access to web-scale training data, increased compliance risk, and the urgent need for standard, enforceable mechanisms that balance research use, copyright, and fair compensation.
Loading comments...
login to comment
loading comments...
no comments yet