Is your client's site blocking AI crawlers? How to check robots.txt for GPTBot, ClaudeBot and PerplexityBot
Which AI crawlers exist, what each one does, how to check whether robots.txt blocks them, and how to decide what a business should allow.
When a customer asks ChatGPT, Claude or Perplexity for "a reliable roofer near me", the answer is built from what those systems can fetch and read. A site that blocks their crawlers cannot be cited, however good its content is. And a surprising number of small business sites block them without anyone deciding to: a security plugin added a blanket rule, a developer copied a robots.txt from a publisher that had opted out, or a hosting platform changed a default.
Checking takes a minute. Deciding what to allow takes a little more thought.
The crawlers that matter
AI companies now run several distinct crawlers, and they do different jobs. The distinction matters because blocking one is not the same as blocking another.
User agent | Operator | What it is for |
|---|---|---|
| OpenAI | Collecting content that may be used to train models |
| OpenAI | Finding and surfacing sites in ChatGPT search results |
| OpenAI | Fetching a page when a user asks ChatGPT to look at it |
| Anthropic | Collecting content for model training |
| Perplexity | Indexing sites to surface in Perplexity answers |
| Controls use of content for Gemini models; does not affect Google Search | |
| Apple | Controls use of content for Apple's AI models |
| Common Crawl | Open web archive widely used for training datasets |
Two details trip people up. First, blocking Google-Extended does not remove a site from Google Search or from AI Overviews, which are built from the normal Googlebot index. Second, a search-oriented crawler such as OAI-SearchBot is separate from the training crawler GPTBot, so a business can opt out of training while still being findable in AI search. Crawler names and purposes change, so check each operator's documentation before advising a client.
How to check
Open https://example.com/robots.txt and look for three patterns.
A blanket block:
User-agent: * Disallow: /This blocks every well-behaved crawler, including search engines. It is usually a staging-site setting that survived launch, and it is an urgent finding on its own.
Named AI blocks:
User-agent: GPTBot Disallow: /Deliberate or not, this stops that crawler. Note which crawlers are named and whether they are the training or the search variety.
Wildcard rules that catch more than intended, such as a plugin that disallows everything except Googlebot and Bingbot.
Also check that robots.txt returns a 200 response. A robots.txt that errors with a 5xx can cause crawlers to back off the whole site.
Remember that robots.txt is a request, not a lock. Reputable crawlers follow it; it does not stop anyone determined to ignore it.
Deciding what to allow
There is no single right answer, and it is the business's decision, not the agency's. A useful way to frame it:
Most local businesses want to be recommended. A plumber, dentist or café gains from appearing in AI answers and loses little from being read. Allowing the search and user-triggered crawlers is usually the right call.
Training is a separate question. Some owners are comfortable with their content being used to train models; some are not. Opting out of training crawlers while allowing search crawlers is a reasonable middle ground.
Publishers and content businesses whose content is the product may reasonably block more.
Whatever the choice, it should be a choice. The finding worth raising is a block nobody meant to put there.
Access is only half of it
Being allowed in does not mean being cited. Answer engines favour content with a clearly identified business entity (structured data such as LocalBusiness or Organization), named authors and visible dates, and pages structured as answers to real questions. We cover that side in auditing a site for AI visibility.
And llms.txt, the proposed file that summarises a site for language models, is worth adding if it is cheap to do, but it is a weak signal compared with robots.txt access and good structure. Sell outcomes, not files.
How to explain it to the owner
When someone asks ChatGPT or Perplexity for a recommendation in your area, your website cannot be one of the sources, because a setting on your site asks those tools not to read it. It looks like this was added by a plugin rather than on purpose. Changing it takes a few minutes.
That lands better than any discussion of user agents.
Check it now
Our free AI robots.txt generator reads a site's current rules, shows which AI crawlers are allowed, and generates a robots.txt and llms.txt for the policy you choose. In a full SiteAssay audit a blocked AI crawler caps the AI visibility score rather than being averaged away, because no amount of good content outweighs being unreadable. The AI visibility page explains how that category is scored, and the website audit checklist covers everything else worth checking.
Written by
SiteAssay