Patreon stops relying on robots.txt alone to curb AI scraping
Patreon is strengthening its use of Cloudflare to block known crawlers that collect creator work for model training. The change is not about closing the door to all automation, but about distinguishing training, search and other uses.
On July 1, 2026, Patreon announced that it was expanding protections against unauthorised AI scraping with Cloudflare. Its stated aim is to block known crawlers that collect creators’ work for model training at the network level, while not blocking search crawlers that can help people discover a page and send visits back to it.
The distinction matters. This is not a move to remove all automation from Patreon, nor does it settle the wider debate about training data. It is a technical and editorial decision about which kinds of access a creator platform wants to permit and under what conditions.
From a request to a technical control
Patreon says it has had measures in place since 2023 to deter AI companies from training on work without permission. The new announcement describes an integration with Cloudflare’s AI Crawl Control to update crawler policies and the tools used to enforce them.
According to Patreon, the change responds to a product reality. Some material that previously sat behind paid membership can now circulate as free media through discovery experiences, including its redesigned Home feed and Quips. A paywall offered de facto protection from some crawlers. As the company opens more content to help creators reach new audiences, it says it wants the protection to travel with that content.
The platform says it will block known training crawlers at the network level. In early testing, it reports that weekly access attempts by individual training crawlers fell from thousands to zero. That is a company claim rather than an independent audit, but it illustrates the intended outcome: not merely declaring a preference, but rejecting identified requests before they reach the content.
Not every bot does the same job
Cloudflare has proposed a more precise classification than the broad label “AI bot.” Its documentation separates three main uses: search, agents and training. A search crawler collects or indexes material to answer queries later; an agent acts in real time on a person’s behalf; and a training crawler collects material to train or fine-tune a model.
The difference affects the balance between control and discovery. Patreon says it will continue to allow search crawlers that can direct readers to creators’ original pages while restricting those dedicated to training. The policy recognises that visibility is not a minor benefit for someone publishing online, yet it does not require accepting every kind of reuse in order to obtain it.
Cloudflare notes that some crawlers can serve more than one purpose. The same operator could index for search and collect data for training. Its system therefore classifies bots by behaviour and lets customers manage separate categories. That is an advance over an all-or-nothing approach, although its effectiveness will depend on operators identifying themselves honestly and on the provider detecting changes in behaviour.
What robots.txt can and cannot do
The Robots Exclusion Protocol is documented by the IETF as a way for a site to indicate rules to crawlers. It is useful for cooperation among services, but it is not authentication or a firewall. Cloudflare’s own options state that preferences published in robots.txt do not directly block access.
That is Patreon’s shift: moving from asking a crawler not to collect certain materials to using a network layer to reject crawlers it identifies as intended for training. No bot list will be perfect. New identifiers can appear, traffic can imitate an ordinary visitor, and disputes can arise over whether automation belongs to search, an agent or training.
Technical blocking therefore does not remove the need for clear rules and transparent relationships with AI operators. But it gives creators something more tangible than a voluntary instruction.
Patreon’s initiative outlines one possible model for the open web: allow indexing that returns audience, restrict collection for model training without permission, and make visible which kind of bot is knocking at the door. The challenge will be applying that separation consistently. For creators, the measure’s value will depend less on its slogan than on whether their work can keep finding an audience without automatically becoming raw material for every model.
This article was produced with artificial intelligence under human editorial oversight.