Git.kernel.org Is "Interesting" to Crawlers (people.kernel.org)

🤖 AI Summary
Git.kernel.org is experiencing significant strain from AI crawlers, which are disproportionately consuming server resources by rendering numerous HTML commits instead of using more efficient methods, like cloning the repositories. Overwhelming traffic from scrapers—estimated at 6 million daily requests, with only 2% being legitimate—ties up around 20% of the site's CPU capacity, hindering its overall performance. The targeted nature of this scraping stems from the rich, unprocessed data available in Linux development, making it an attractive source for training large language models (LLMs). The situation highlights a persistent challenge in the AI/ML community: the exploitation of public repositories for data collection while bypassing conventional data use protocols. Efforts to mitigate this influx, such as implementing difficulty-based challenges for bots, have proven only temporarily effective, as many are now capable of overcoming these barriers. As a response, Git.kernel.org plans to limit access and features for anonymous users, an action that balances ongoing accessibility with the necessity of protecting their infrastructure from excessive scraping. This scenario underscores the need for better data usage practices and the consideration of impacts on shared resources in the AI training landscape.
Loading comments...
loading comments...