Obituary: Farewell to robots.txt (1994-2025) (www.heise.de)

🤖 AI Summary
robots.txt — the 1994 convention-born text file formalized as RFC 9309 in 2022 — has effectively been retired as a universal traffic cop for the web. Decades of voluntary compliance unraveled under the scale of modern AI crawlers: the web’s old crawl-to-referral economics (Google once averaged ~14:1) broke down when AI bots generated massive fetches with almost no downstream referrals. Cloudflare’s July 2025 decision to block AI crawlers by default symbolized the collapse of voluntary norms; measurements showed industrial-scale extraction (OpenAI-linked crawls hitting ~1,700:1 crawl-to-referral, Anthropic’s ClaudeBot in the tens of thousands to one). The Internet Archive’s 2017 choice to ignore robots.txt and tactics like undisclosed IPs by Perplexity further signaled that polite opt-outs no longer constrain large actors. Technically and socially, the loss matters because it shifts the governance model from a lightweight, interoperable signal to an arms race of enforcement and fragmentation. Site operators are deploying TLS-based crawler fingerprinting, honeypots and behavioral analysis while regulators (EDPB Opinion 28/2024, Italy’s €15M OpenAI fine) push legal remedies. Proposed successors (ai.txt, TDM ReP, No-AI-Training headers) compete with commercial paywalls, licensing deals and network-level blocks, risking a fractured web where access depends on contracts or technical gatekeeping. For the AI/ML community this means harder, messier access to web-scale training data, increased compliance risk, and the urgent need for standard, enforceable mechanisms that balance research use, copyright, and fair compensation.
Loading comments...
loading comments...