Block, allow or price: the AI crawler decision after Cloudflare's September flip
Cloudflare's 15 September 2026 default is a publisher story. The decision a normal business site has to make is narrower: which crawlers to block, which to allow, and whether pricing is even available to you.
Short answer
For a business site with no ad slots, allow the answering crawlers (OAI-SearchBot, Claude-SearchBot, PerplexityBot, Bingbot), treat the training crawlers (GPTBot, ClaudeBot, Google-Extended) as a separate and lower-stakes decision, and skip pricing entirely. Blocking an answering crawler removes you from that surface. Blocking a training crawler does not. Cloudflare's new default targets pages that display ads, which a clinic, restaurant or agency site does not have.
On 15 September 2026 Cloudflare flipped a default, and every recap you read that week was written for publishers.
Here is the scope.
The new posture blocks mixed-use AI crawlers by default on pages that display ads, and TechCrunch's 1 July 2026 piece put the customer classes at new Cloudflare customers, new sites set up by existing customers, and all existing free customers. On the pricing side, Pay Per Crawl, launched in Cloudflare's own announcement, is becoming Pay Per Use.
That is a negotiation between large content owners and large model labs.
You are running a clinic, a restaurant, a studio or a twelve-page services site with no ad slots anywhere on it. So the flip mostly has nothing to act on, and the decision you still have to make is the one nobody wrote up.
We nearly shipped the wrong version of it. The draft of this tree sitting in the NovaTechRay docs in July said block Training, allow Search, because that reads as the cautious middle path. Then we read Cloudflare's own definition of Training, which folds in mixed-purpose crawlers used for both Training and Search, and deleted the branch.
Cautious was the wrong instinct. There are only three moves available to you, and two of them are usually wrong.
Branch 1: does the flip reach you at all
Two conditions, both required.
You have to sit inside the customer classes TechCrunch listed. And the page has to display ads, because that is the trigger Cloudflare scoped the new default to.
Most business sites fail the second condition outright.
A menu page, a services page and a contact page carry no ad inventory, so there is nothing for an ad-page default to apply to.
If both conditions miss you, the September date is not your problem and you can stop reading the news coverage. What you still have is a settings panel you have probably never opened. We wrote the diagnostic for that separately in is Cloudflare blocking ChatGPT from your website, and it is worth ten minutes even if the answer turns out to be no.
Branch 2: the answering crawlers, which you allow
This is the branch with a real cost attached, and it is the one people get backwards.
OAI-SearchBot, Claude-SearchBot, PerplexityBot and Bingbot are answering crawlers.
They fetch your page so it can be quoted, summarised and cited inside a live answer, which is the entire mechanism by which a stranger asking about your category ends up reading your name.
Block one and you are removed from that surface. Not demoted. Removed.
There is no partial credit, no residual presence, no slow decay you can reverse next quarter.
The engine simply has nothing of yours to retrieve, so it answers using whoever did not block it.
But the reason people block them anyway is that the bot names look like the training bots, and the dashboard groups them under one scary heading. A row that says "AI crawler" invites a single toggle. The toggle is where the damage happens.
So the rule here has no exceptions worth the paragraph it would take to write. Allow all four. If you want the per-directive version at the robots.txt layer rather than the edge layer, the AI crawlers and robots.txt guide has it.
Branch 3: the training crawlers, which are a genuine choice
GPTBot, ClaudeBot and Google-Extended are a different category and a much smaller decision.
Blocking GPTBot does not remove you from ChatGPT. Blocking ClaudeBot does not remove you from Claude.
And Google-Extended is not even a crawler, it is a token controlling whether already-crawled content may be used for Gemini and for grounding, and disallowing it leaves ordinary Google Search untouched.
Which means the training decision is a licensing preference, not a visibility one.
For a hospitality or services site, there is usually nothing on the page worth withholding. Your opening hours are not intellectual property. If you run a genuine content asset, a research library, a course, a set of original guides, blocking training is defensible and costs you nothing on the answer surfaces.
The trap is the mixed-purpose clause. Cloudflare's Training definition explicitly covers crawlers used for both Training and Search, so a blunt Training block can reach further than you intended. Read the category definition in your own dashboard before you flip it, because the words Search and Training are doing specific work there and not the plain-English work you would assume.
Branch 4: pricing, which is not available to you
Pay Per Crawl was announced as a marketplace.
The publisher sets a per-request price, and a crawler that has not paid gets refused rather than served, which Cloudflare's Pay Per Crawl post frames as giving content owners a lever they previously did not have. Pay Per Use is the same lever with a broader name.
A marketplace needs a buyer on the other side.
Nobody is bidding for a twelve-page dental site. There is no scarcity, no archive, no unique corpus, so the price you set is not a negotiation, it is a refusal with extra steps. You end up in the same place as Branch 2, with the answering crawlers turned away, except you also believe you are monetising something.
This is the branch where the publisher recap and the business site part company completely.
Pricing is a real strategy if your content is the product. If your content is marketing for a product, pricing it is a way of hiding your own marketing.
Branch 5: what to check after you decide
Whatever you choose, the setting and the outcome are different things and you should confirm both.
Read the robots.txt your edge serves rather than the one in your repository, because a managed file can prepend rules you never wrote.
Then check that a plain unauthenticated fetch of your homepage returns your content and not a challenge page.
Then open the crawler log. Cloudflare's AI Crawl Control logs which AI services requested your content, and that log is the only evidence in this whole exercise. Everything else, including this article, is inference.
One honest limit. A curl test with a pasted user-agent string does not reproduce what a verified crawler experiences, because verification runs on signature and network rather than on the string you typed. So a clean curl is a weak pass, not a certificate.
We run this check on every NovaTechRay engagement before we write a word of copy, for a boring reason.
And content work on a site that refuses the crawler is spend with no delivery mechanism attached.
The short version
- Check whether the September default reaches you at all. No ad slots usually means no.
- Allow OAI-SearchBot, Claude-SearchBot, PerplexityBot and Bingbot, always.
- Treat GPTBot, ClaudeBot and Google-Extended as a separate licensing call, and read Cloudflare's Training definition before flipping it.
- Skip Pay Per Crawl unless your content is the product rather than the advertisement for it.
- Verify against the edge-served robots.txt and a real fetch, not against your repository.
- Read the crawler log afterwards and treat it as the only evidence.
If the branch you land on is unclear, that is the kind of thing NovaTechRay untangles in an hour at novatechray.com.
Frequently asked questions
Should a small business block AI crawlers on Cloudflare?
Almost never the answering ones. OAI-SearchBot, Claude-SearchBot, PerplexityBot and Bingbot are how your pages get fetched and cited while a customer waits for an answer, so blocking them removes you from that surface with nothing gained. Training crawlers are a separate call and a lower-stakes one, because blocking GPTBot or ClaudeBot does not remove you from the corresponding answer engine.
Can I charge AI crawlers to access my site?
Technically yes, practically no, if you are a business rather than a publisher. Cloudflare's Pay Per Crawl announcement set it up as a marketplace where the publisher names a per-request price and an unpaid crawler gets refused. A marketplace needs a buyer. Nobody is bidding for a twelve-page dental site, so the only outcome of pricing it is the refusal.
Does the 15 September 2026 default apply to my site?
Only if two things are true: you fall inside the customer classes TechCrunch listed on 1 July 2026, which were new customers, new sites from existing customers, and all existing free customers, and the page in question displays ads. A business site with no ad slots gives the rule nothing to act on.
What is the difference between blocking a training crawler and blocking an answering crawler?
A training crawler collects content that may end up inside a future model, with no attribution and no traffic. An answering crawler fetches your page so it can be quoted and cited in a live answer. Blocking the first is a licensing preference. Blocking the second is a visibility decision that takes you off that engine.
Want to know how AI models currently describe your business?
We run a free visibility check across ChatGPT, Perplexity, Claude and Google AI Overviews, then show you exactly which signals are missing.
Book a visibility check