Cloudflare’s new crawler defaults take effect today, September 15, 2026. New Cloudflare customers, new sites added by existing customers, and every free-plan site that has not changed its dashboard settings will now allow AI crawlers to index pages for search, while blocking those same pages from being used for model training and agent retrieval if the page carries ads. Crawlers that refuse to separate search from training and agent use are blocked outright on ad-supported pages.
The deadline was set on July 1, when Cloudflare gave the AI industry a two month window to split its bots by purpose. The company confirmed the date in its own announcement, Cloudflare Allows the Agentic Internet to Flourish with a Simple Philosophy: Your Content, Your Rules, and site owners can still override the defaults from their dashboard at any time.
What counts as a mixed-use crawler
A transparent AI company runs separate bots with separate names, one for search indexing, one for training data collection, one for live agent retrieval. A site owner can then allow one and refuse another. A mixed-use crawler collapses all three into a single bot, so saying yes to search discovery also means saying yes to training.
That is the behaviour Cloudflare is targeting. In the company’s framing, mixed crawlers force publishers into a choice between staying findable and giving away their most valuable work for nothing. Matthew Prince, Cloudflare’s co-founder and CEO, said in the announcement that the company hopes the default changes “encourage mixed use crawlers to separate out search from agent use and training.”
The numbers behind the argument
Cloudflare states that automated agents and bots now drive more than half of all web requests. It also says the largest search engine currently has access to roughly twice as much information as leading AI companies, precisely because it is hard for a site to stay discoverable there without also feeding AI systems.
A third figure is the one publishers will feel in their bandwidth bills. Cloudflare’s data suggests that over 50 percent of crawl traffic from AI crawlers is spent re-fetching pages that have not changed. The company is testing freshness signals with AI firms to cut that waste, and says it plans to make them broadly available later this year.
On the money side, Cloudflare counted more than 50 major content licensing agreements signed between publishers and AI platforms over the past year, and has moved its Pay Per Crawl product to a Pay Per Use model, where a publisher is paid when their content is actually used in an answer rather than when a bot merely fetches it. Ceramic.ai and You.com are the first partners. Details are in Cloudflare’s blog post Your site, your rules: new AI traffic options for all customers.
Who is affected, and who is not
Existing paid Cloudflare customers keep whatever settings they already have. Nothing changes for them automatically. The change applies to new customers, newly added sites, and free-plan users who have not touched the AI settings in their dashboard.
That last group is the important one for readers here. A large share of small business sites, personal blogs, portfolio sites and client projects across Pakistan and the Gulf sit on Cloudflare’s free plan, often because a developer added it years ago for speed and DDoS protection and then never opened the dashboard again. Those sites just had a meaningful content policy applied on their behalf.
Three things to check on your own sites this week
First, log into the Cloudflare dashboard for every domain you manage, including client domains, and look at the AI crawler settings rather than assuming the default suits you. Second, decide deliberately whether you want AI training access blocked. If you are a portfolio site or a lead-generation site with no ads, being quoted by an answer engine is usually free marketing, and you may want the opposite of the new default. Third, if you do run ads, check whether your traffic mix actually depends on AI referrals before you leave the block in place.
The wider shift is the same one we covered when Google’s AI search started changing how people use the web. Discovery is moving from a list of links to a generated answer, and the question for anyone who publishes online is no longer only how to rank, but whether being used at all is worth what it costs you.
The part nobody can predict yet
Cloudflare sits in front of a very large slice of the web, which makes a default setting closer to a policy than a preference. But the outcome depends on a response the company cannot control. If mixed-use crawlers split their bots by purpose, the standard spreads and publishers get real choice. If they do not, a meaningful chunk of the open web becomes invisible to AI systems that many readers now use as their first stop. Which of those two happens will probably be clear within a few months of crawl data, not years.




