One domain, one card: whether the crawlers behind ChatGPT, Claude, Perplexity and Gemini may read your site, whether it serves an llms.txt, and where its XML sitemap is. Every verdict shows the rule that decided it, and the robots.txt lines to paste come with the result.
18 checks, three files. robots.txt, with a verdict for the site root (/) for 14 AI crawlers and Googlebot; /llms.txt; and an XML sitemap, looked for in the robots.txt Sitemap: lines first, then four usual paths. As VerandBot/1.0, redirects followed, 10 seconds per fetch. A non-200 answer or an HTML page at /robots.txt counts as no file.
The crawler's own group if one names it, the longest matching name winning, else the * group. Inside it, the longest rule that matches / decides, and on a tie Allow wins. No group at all means allowed. The same rule decides all 15, so every crawler named on the card gets its own verdict and the line behind it.
Only the root is judged, so Disallow: /blog/ still reads as allowed here. Not seen: meta robots tags, X-Robots-Tag headers, firewall or CDN bot rules, text that only JavaScript renders. The llms.txt and sitemap contents are not validated. Whether a crawler obeys the file is its operator's call, and OpenAI and Perplexity both say their user-triggered fetchers may not.
Three files at the root, one verdict per crawler, one fix. The card is the tool in motion on an example domain, looped, and each step lights up while the card is doing it.
User-agent: GPTBot Allow: /
You type a host, not a page. The tool asks that host for /robots.txt and /llms.txt as VerandBot, and then goes looking for a sitemap. Each fetch follows redirects and gives up after ten seconds.
An error, a timeout or an HTML page at a text path counts as no file. The sitemap is tried where robots.txt declares it, then at four usual paths, and only XML that parses as a sitemap counts.
For each AI crawler, and Googlebot for contrast: which group speaks to it, and which line decides whether it may fetch the root. The card prints both, so a surprise can be traced to one line.
Every result carries allow groups for ten AI crawlers: the assistants' own, not the four bulk scrapers, which stay your call. Each group repeats your * rules, so what you block stays blocked. A crawler you already name is left as written, so edit its own group to change it; paste only the crawlers you mean to let in.
Whether an assistant can cite you is decided per host, before any page is read, by a file that is easy to forget exists. Here is how the decision is made, what sits beside it, and how to set a policy on purpose rather than by accident.
A robots.txt checker, or a robots.txt validator, answers a narrow question: is this file well formed, and what does each line mean. That job belongs to our Robots.txt Checker, which reads its groups and rules. This tool asks the question a firm actually has: taken together, what does my site tell AI crawlers? Three files at the root of the host carry that answer. robots.txt says who may fetch what. The XML sitemap says what exists. llms.txt, where there is one, is a curated list of the pages you would point a reader to first.
The unit is the host, and that catches people out. Crawlers fetch /robots.txt from each host separately, so help.yourdomain.com and www.yourdomain.com each need their own files, and a rule on the main site says nothing about the help centre. Our own runs show it: on the day this page was built, verand.ai had no robots.txt, no llms.txt and no sitemap, while help.verand.ai, a separate host, served a 17-URL sitemap. Two hosts, one company, two different answers. Check each host people land on.
The single most useful thing to know about AI crawlers is that each major operator now runs several, and they do different jobs. One collects pages for training future models. One builds the index an assistant searches when it answers. One fetches a page in the moment because a person asked a question that needs it. Blocking one does not block the others, and the consequences are not the same. Here is how each operator documents its crawlers today:
| Operator | Training | Search index | Fetch for a user |
|---|---|---|---|
| OpenAI | GPTBot | OAI-SearchBot | ChatGPT-User |
| Anthropic | ClaudeBot | Claude-SearchBot | Claude-User |
| Perplexity | none named | PerplexityBot | Perplexity-User |
| Google-Extended (a token) | Googlebot | not separately named |
A few lines from those documents matter more than the rest. OpenAI says disallowing GPTBot does not affect ChatGPT search, which uses OAI-SearchBot, and that because ChatGPT-User acts for a user, "robots.txt rules may not apply". Perplexity says PerplexityBot "is not used to crawl content for AI foundation models", and that Perplexity-User "generally ignores robots.txt rules" because a user requested the fetch. Google says Google-Extended "does not impact a site's inclusion in Google Search nor is it used as a ranking signal". Anthropic lists ClaudeBot for training data, Claude-User for fetches when someone asks Claude a question, and Claude-SearchBot for search, and says all three honour robots.txt. The card gives each of the six its own verdict.
The card does not just look for a scary line. It reads the file the way Google's open-source parser does, and prints the evidence. First the group: the one whose User-agent name matches the crawler, the longest name winning, and only if there is none, the * group. Then the rule: inside that group, of every Allow and Disallow path that matches the root, the longest decides, and a tie goes to Allow. That is why a file that looks selective can do the opposite of what it seems to say:
# looks like "AI search in, training out"
User-agent: *
Disallow: /
User-agent: OAI-SearchBot
User-agent: PerplexityBot
Allow: /
User-agent: GPTBot
Disallow: /client-portal/
OAI-SearchBot and PerplexityBot are allowed: their own group says Allow: /. GPTBot is allowed too, which is probably not what was meant: its own group applies, and /client-portal/ does not match the root. ClaudeBot has no group, falls to *, and is blocked. So is Googlebot, which means this file takes the site out of Google Search as well. Every one of those outcomes is on the card, with the group and the rule beside it.
One edge to know. RFC 9309 says that when two groups name the same crawler, their rules are combined; this checker reads the first group that names it. Keep one group per crawler and both readings agree, which is also the only way a human reviewing the file will read it correctly.
An allowed verdict means the file does not stop a crawler that honours it. It does not mean the crawler gets the page. Several things sit after robots.txt, and none of them is visible to a file check. A CDN or security plugin can refuse AI crawlers at the network edge, and some offer exactly that as a switch, so the file says yes while the firewall says no. A page can carry a noindex meta tag or an X-Robots-Tag header. Text that only appears after JavaScript runs may never reach a crawler that reads the served HTML. Other checkers make this point well too. The practical step is the same: when the card says allowed and an assistant still cannot see you, ask whoever runs your CDN or firewall what it does with the names on this card, and test a single URL with the AI Crawler Access Checker.
llms.txt is a proposed convention from 2024: a Markdown file at /llms.txt listing the pages a site considers its best summary of itself. It is not a standard, and no crawler is obliged to fetch it. The card reports it as present or missing, and shows a missing file as a warning rather than a failure, because it blocks nothing. For a regulated firm its real value is internal: a short list of pages that have been reviewed and that you are content to have quoted. The llms.txt Generator drafts one from your pages.
The sitemap is found the way crawlers find it: the Sitemap: lines in robots.txt first, then /sitemap.xml, /sitemap_index.xml, /wp-sitemap.xml and /sitemap-index.xml. The card says where it was found, how many URLs it lists, and whether robots.txt declares it. Willowdale Equity's file declares its sitemap and the sitemap lists 269 URLs; help.verand.ai serves one at the usual path with no robots.txt to declare it in. Whether the XML itself is sound is a separate question, answered by the Sitemap Checker and the Sitemap Validator.
Most AI crawler blocks were never decided. They arrived with a hosting panel, a security plugin, a staging file that shipped, or a quick reaction to a news story. For a firm that publishes expertise so clients can find it, a better approach is three decisions, made once and written down:
For the firm that wants to be cited but keep its writing out of training sets, the file looks like this. It is written for this page, not returned by the tool, and the checker will show GPTBot, ClaudeBot and Google-Extended as blocked when it is in place, which is then the intended result:
# training crawlers: withheld on purpose User-agent: GPTBot User-agent: ClaudeBot User-agent: anthropic-ai User-agent: Google-Extended User-agent: CCBot Disallow: / # search and user fetch: allowed so answers can cite us User-agent: OAI-SearchBot User-agent: ChatGPT-User User-agent: Claude-SearchBot User-agent: Claude-User User-agent: PerplexityBot User-agent: Perplexity-User Allow: / # everyone else User-agent: * Disallow: /client-portal/ Sitemap: https://yourdomain.com/sitemap.xml
For the firm that is content to be read by all of them, the allow snippet on the card is the whole answer. For a site with no robots.txt at all, nothing is blocked today, but nothing is recorded either: the snippet plus a Sitemap: line is a sound first file.
* applies to it, including the paths you meant to keep private.Sitemap: line, crawlers that probe the usual paths will not find it./llms.txt is not a file, and this checker reports it as missing.Six things that are true of this tool, each one backed by a line in the code that runs it.
The request carries a domain and nothing else. There is no account, no session and no database behind the tool, so there is nothing for us to keep about you.
robots.txt, llms.txt and the sitemap come back from one request, so the answer to "what does my site tell AI crawlers" is on one screen rather than across three tools.
Each crawler's row names the group that applied and the line that decided it, parsed by a fixed grammar with no model involved. The same file gives the same answer every time.
The robots.txt, llms.txt, sitemap and blocked-crawler checks are the site-level checks Verand's deep crawl runs on every customer site every two weeks, called directly. The per-crawler verdicts add a finer parser on top.
Only the root is judged, and firewalls, headers, meta tags and rendered text are out of reach. The card says so beside the result rather than in a footnote.
A run is a handful of small text fetches, so it costs nothing and is never metered. The one limit is a courtesy to the sites being fetched: 20 checks a minute per visitor.
Which crawlers to allow, what blocking one really does, and where robots.txt stops.
Split them by job. Allow the search crawlers and user fetchers, OAI-SearchBot, ChatGPT-User, PerplexityBot, Perplexity-User, Claude-SearchBot and Claude-User, because they are how an assistant finds and cites a page when someone asks. Make a deliberate call on the training crawlers, GPTBot, ClaudeBot and the Google-Extended token: allowing them suits educational content you want widely known, and blocking them suits writing that is itself the product or is covered by a client agreement. If you do block AI crawlers, block those names specifically rather than with a Disallow: / under *, which also shuts out Google Search. This is a content decision, not legal advice.
A training bot collects pages that may be used to train future models; GPTBot and ClaudeBot are documented that way. A search bot builds the index an assistant searches when it answers, as OAI-SearchBot does for ChatGPT search and PerplexityBot does for Perplexity. A user-fetch bot, such as ChatGPT-User or Perplexity-User, requests one page because a person asked a question that needs it. OpenAI says robots.txt rules may not apply to ChatGPT-User, and Perplexity says Perplexity-User generally ignores them, because a user requested the fetch.
No. OpenAI documents GPTBot as its training crawler and says disallowing it does not affect ChatGPT search, which uses OAI-SearchBot; live fetches for a user come from ChatGPT-User. So a GPTBot block withholds your pages from training and leaves you eligible to be found and cited. Blocking OAI-SearchBot is what keeps you out of ChatGPT search answers, which OpenAI says may still show your site as a navigational link.
robots.txt is only the first gate. A CDN, web application firewall or security plugin can refuse a crawler by name or by IP before it ever reaches a page, and some offer AI-bot blocking as a setting, so the file can say allowed while the network says no. A page can also carry a noindex tag or an X-Robots-Tag header, or put its text behind JavaScript. This checker reads the file and cannot see any of that, so when the card says allowed and an assistant still cannot see you, the CDN and firewall settings are the next place to look.
No. Google describes Google-Extended as a product token that controls whether content it crawls may be used to train future Gemini models, and says it "does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search". It is not a separate crawler; Googlebot does the fetching. Blocking it therefore has no effect on rankings, and the only robots.txt line that takes you out of Google Search is one that reaches Googlebot.
Often enough to re-read your file once in a while. Operators add crawlers as their products change: Anthropic now documents Claude-User and Claude-SearchBot alongside ClaudeBot, and anthropic-ai survives mainly as an older token in existing files. This checker tests 14 AI names, the same list Verand's site audit watches, including Claude-User and Claude-SearchBot, plus Googlebot for contrast. A misspelled or outdated name in your file blocks or allows nothing, so check each name against the operator's own documentation.
Access is the first gate, not the finish. Verand writes articles from your own expertise and credentials, so both Google and the AI assistants have something of yours to name. It then tracks where you rank and where ChatGPT, Gemini, Google AI Overviews, Google AI Mode, Perplexity and Claude mention you, and gates every draft so a claim your regulator would not allow never publishes.
Content built to rank in
Google and get cited by
ChatGPT
Perplexity
Gemini
Claude, with every claim checked before it goes live.