HELLFIRE https://www.theregister.com/2025/04/03/wikimedia_foundation_bemoans_bot_bandwidth/ https://diff.wikimedia.org/2025/04/01/how-crawlers-impact-the-operations-of-the-wikimedia-projects/ Web-scraping bots have become an unsupportable burden for the Wikimedia community due to their insatiable appetite for online content to train AI models. Representatives from the Wikimedia Foundation, which oversees Wikipedia and similar community-based projects, say that since January 2024, the bandwidth spent serving requests for multimedia files has increased by 50 percent. “This increase is not coming from human readers, but largely from automated programs that scrape the Wikimedia Commons image catalog of openly licensed images to feed images to AI models,” explained Birgit Mueller, Chris Danis, and Giuseppe Lavagetto, from the Wikimedia Foundation in a public post. … The heedlessness of ill-behaved bots has been a common complaint over the past year or so among those operating computing infrastructure for open source projects, as the Wikimedians themselves noted by pointing to our recent report on the matter. … But since ChatGPT came online and generative AI took off, bots have become more willing to stripmine entire websites for content that’s used to train AI models. And these models may end up as commercial competitors, offering the aggregate knowledge they’ve gathered for a subscription fee or for free. Either scenario has the potential to reduce the need for the source website or for search queries that generate online ad revenue. … Regards
Somewhat-Reticent Rook … but isn’t interested in paying for the content … Perhaps such scraping needs to be metered? More overhead for detection … ow. Robots.txt not honored → tox?