Free AI Visibility Tool · No Signup Required

Free AI Crawler Checker

One domain, one card: whether the crawlers behind ChatGPT, Claude, Perplexity and Gemini may read your site, whether it serves an llms.txt, and where its XML sitemap is. Every verdict shows the rule that decided it, and the robots.txt lines to paste come with the result.

https://

Read-only. A few small fetches from our server, robots.txt, llms.txt and the sitemap paths, and nothing is written to your site.

  • No signup, no email wall
  • 14 AI crawlers by name
  • llms.txt and sitemap too
  • Rule shown per crawler
  • Free, no daily cap
  • Fix ready to paste
What we checked

18 checks, three files. robots.txt, with a verdict for the site root (/) for 14 AI crawlers and Googlebot; /llms.txt; and an XML sitemap, looked for in the robots.txt Sitemap: lines first, then four usual paths. As VerandBot/1.0, redirects followed, 10 seconds per fetch. A non-200 answer or an HTML page at /robots.txt counts as no file.

How a verdict is decided

The crawler's own group if one names it, the longest matching name winning, else the * group. Inside it, the longest rule that matches / decides, and on a tie Allow wins. No group at all means allowed. The same rule decides all 15, so every crawler named on the card gets its own verdict and the line behind it.

What it cannot see

Only the root is judged, so Disallow: /blog/ still reads as allowed here. Not seen: meta robots tags, X-Robots-Tag headers, firewall or CDN bot rules, text that only JavaScript renders. The llms.txt and sitemap contents are not validated. Whether a crawler obeys the file is its operator's call, and OpenAI and Perplexity both say their user-triggered fetchers may not.

About this tool

How the AI Crawler Checker reads a domain.

Three files at the root, one verdict per crawler, one fix. The card is the tool in motion on an example domain, looped, and each step lights up while the card is doing it.

01

One domain, three root files

You type a host, not a page. The tool asks that host for /robots.txt and /llms.txt as VerandBot, and then goes looking for a sitemap. Each fetch follows redirects and gives up after ten seconds.

02

Each file answered, or named missing

An error, a timeout or an HTML page at a text path counts as no file. The sitemap is tried where robots.txt declares it, then at four usual paths, and only XML that parses as a sitemap counts.

03

A verdict per crawler, with its rule

For each AI crawler, and Googlebot for contrast: which group speaks to it, and which line decides whether it may fetch the root. The card prints both, so a surprise can be traced to one line.

04

The fix is a paste

Every result carries allow groups for ten AI crawlers: the assistants' own, not the four bulk scrapers, which stay your call. Each group repeats your * rules, so what you block stays blocked. A crawler you already name is left as written, so edit its own group to change it; paste only the crawlers you mean to let in.

AI crawler access, explained

What AI crawlers read at your root, and how that differs from a robots.txt checker.

Whether an assistant can cite you is decided per host, before any page is read, by a file that is easy to forget exists. Here is how the decision is made, what sits beside it, and how to set a policy on purpose rather than by accident.

Three files at the root, read as one answer

A robots.txt checker, or a robots.txt validator, answers a narrow question: is this file well formed, and what does each line mean. That job belongs to our Robots.txt Checker, which reads its groups and rules. This tool asks the question a firm actually has: taken together, what does my site tell AI crawlers? Three files at the root of the host carry that answer. robots.txt says who may fetch what. The XML sitemap says what exists. llms.txt, where there is one, is a curated list of the pages you would point a reader to first.

The unit is the host, and that catches people out. Crawlers fetch /robots.txt from each host separately, so help.yourdomain.com and www.yourdomain.com each need their own files, and a rule on the main site says nothing about the help centre. Our own runs show it: on the day this page was built, verand.ai had no robots.txt, no llms.txt and no sitemap, while help.verand.ai, a separate host, served a 17-URL sitemap. Two hosts, one company, two different answers. Check each host people land on.

Training, search and user fetch: three different jobs

The single most useful thing to know about AI crawlers is that each major operator now runs several, and they do different jobs. One collects pages for training future models. One builds the index an assistant searches when it answers. One fetches a page in the moment because a person asked a question that needs it. Blocking one does not block the others, and the consequences are not the same. Here is how each operator documents its crawlers today:

OperatorTrainingSearch indexFetch for a user
OpenAIGPTBotOAI-SearchBotChatGPT-User
AnthropicClaudeBotClaude-SearchBotClaude-User
Perplexitynone namedPerplexityBotPerplexity-User
GoogleGoogle-Extended (a token)Googlebotnot separately named

A few lines from those documents matter more than the rest. OpenAI says disallowing GPTBot does not affect ChatGPT search, which uses OAI-SearchBot, and that because ChatGPT-User acts for a user, "robots.txt rules may not apply". Perplexity says PerplexityBot "is not used to crawl content for AI foundation models", and that Perplexity-User "generally ignores robots.txt rules" because a user requested the fetch. Google says Google-Extended "does not impact a site's inclusion in Google Search nor is it used as a ranking signal". Anthropic lists ClaudeBot for training data, Claude-User for fetches when someone asks Claude a question, and Claude-SearchBot for search, and says all three honour robots.txt. The card gives each of the six its own verdict.

How a verdict is reached: groups and precedence

The card does not just look for a scary line. It reads the file the way Google's open-source parser does, and prints the evidence. First the group: the one whose User-agent name matches the crawler, the longest name winning, and only if there is none, the * group. Then the rule: inside that group, of every Allow and Disallow path that matches the root, the longest decides, and a tie goes to Allow. That is why a file that looks selective can do the opposite of what it seems to say:

# looks like "AI search in, training out"
User-agent: *
Disallow: /

User-agent: OAI-SearchBot
User-agent: PerplexityBot
Allow: /

User-agent: GPTBot
Disallow: /client-portal/

OAI-SearchBot and PerplexityBot are allowed: their own group says Allow: /. GPTBot is allowed too, which is probably not what was meant: its own group applies, and /client-portal/ does not match the root. ClaudeBot has no group, falls to *, and is blocked. So is Googlebot, which means this file takes the site out of Google Search as well. Every one of those outcomes is on the card, with the group and the rule beside it.

One edge to know. RFC 9309 says that when two groups name the same crawler, their rules are combined; this checker reads the first group that names it. Keep one group per crawler and both readings agree, which is also the only way a human reviewing the file will read it correctly.

robots.txt is the first gate, not the only one

An allowed verdict means the file does not stop a crawler that honours it. It does not mean the crawler gets the page. Several things sit after robots.txt, and none of them is visible to a file check. A CDN or security plugin can refuse AI crawlers at the network edge, and some offer exactly that as a switch, so the file says yes while the firewall says no. A page can carry a noindex meta tag or an X-Robots-Tag header. Text that only appears after JavaScript runs may never reach a crawler that reads the served HTML. Other checkers make this point well too. The practical step is the same: when the card says allowed and an assistant still cannot see you, ask whoever runs your CDN or firewall what it does with the names on this card, and test a single URL with the AI Crawler Access Checker.

llms.txt and the sitemap: what the card reports

llms.txt is a proposed convention from 2024: a Markdown file at /llms.txt listing the pages a site considers its best summary of itself. It is not a standard, and no crawler is obliged to fetch it. The card reports it as present or missing, and shows a missing file as a warning rather than a failure, because it blocks nothing. For a regulated firm its real value is internal: a short list of pages that have been reviewed and that you are content to have quoted. The llms.txt Generator drafts one from your pages.

The sitemap is found the way crawlers find it: the Sitemap: lines in robots.txt first, then /sitemap.xml, /sitemap_index.xml, /wp-sitemap.xml and /sitemap-index.xml. The card says where it was found, how many URLs it lists, and whether robots.txt declares it. Willowdale Equity's file declares its sitemap and the sitemap lists 269 URLs; help.verand.ai serves one at the usual path with no robots.txt to declare it in. Whether the XML itself is sound is a separate question, answered by the Sitemap Checker and the Sitemap Validator.

A policy for a regulated firm, not a default

Most AI crawler blocks were never decided. They arrived with a hosting panel, a security plugin, a staging file that shipped, or a quick reaction to a news story. For a firm that publishes expertise so clients can find it, a better approach is three decisions, made once and written down:

  • Allow search and user fetch. OAI-SearchBot, PerplexityBot, Claude-SearchBot and the user fetchers are how an assistant finds and quotes a page when someone asks. Blocking the search crawlers keeps you out of those answers without protecting anything a training crawler would take.
  • Decide on training deliberately. Letting GPTBot, ClaudeBot and Google-Extended read your articles is a reasonable choice for educational content you want widely known. Withholding them is equally reasonable where the writing is the product, or a client agreement or licence says so. Either way it is a content decision, not legal advice; your counsel decides what an agreement requires.
  • Confirm nothing overrides the file. Ask IT or the agency that runs the CDN whether any bot rule applies to these names. Then re-check this card after every deploy, because robots.txt is the file that gets overwritten.

For the firm that wants to be cited but keep its writing out of training sets, the file looks like this. It is written for this page, not returned by the tool, and the checker will show GPTBot, ClaudeBot and Google-Extended as blocked when it is in place, which is then the intended result:

# training crawlers: withheld on purpose
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: anthropic-ai
User-agent: Google-Extended
User-agent: CCBot
Disallow: /

# search and user fetch: allowed so answers can cite us
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
Allow: /

# everyone else
User-agent: *
Disallow: /client-portal/

Sitemap: https://yourdomain.com/sitemap.xml

For the firm that is content to be read by all of them, the allow snippet on the card is the whole answer. For a site with no robots.txt at all, nothing is blocked today, but nothing is recorded either: the snippet plus a Sitemap: line is a sound first file.

Mistakes that only show at the site level

  • A subdomain with no files of its own. The docs, help centre or client portal on its own host inherits nothing from the main site's robots.txt.
  • www and the bare domain serving different files. Check the host your links and sitemap actually use.
  • A named group that quietly drops the general rules. Once a crawler has its own group, nothing in * applies to it, including the paths you meant to keep private.
  • Blocking the user fetchers to stop training. It leaves the training crawlers exactly where they were. Claude-User honours the block and stops fetching for Claude's users; OpenAI and Perplexity say their user fetchers may not follow robots.txt at all.
  • A sitemap only at an unusual path. Without a Sitemap: line, crawlers that probe the usual paths will not find it.
  • llms.txt answering with an HTML page. A theme's 404 page served at /llms.txt is not a file, and this checker reports it as missing.
Why this one

Why choose Verand's AI Crawler Checker?

Six things that are true of this tool, each one backed by a line in the code that runs it.

No signup, no email wall

The request carries a domain and nothing else. There is no account, no session and no database behind the tool, so there is nothing for us to keep about you.

Three files in one card

robots.txt, llms.txt and the sitemap come back from one request, so the answer to "what does my site tell AI crawlers" is on one screen rather than across three tools.

The rule behind every verdict

Each crawler's row names the group that applied and the line that decided it, parsed by a fixed grammar with no model involved. The same file gives the same answer every time.

The product's own checks

The robots.txt, llms.txt, sitemap and blocked-crawler checks are the site-level checks Verand's deep crawl runs on every customer site every two weeks, called directly. The per-crawler verdicts add a finer parser on top.

Names what it cannot see

Only the root is judged, and firewalls, headers, meta tags and rendered text are out of reach. The card says so beside the result rather than in a footnote.

$0, no daily cap

A run is a handful of small text fetches, so it costs nothing and is never metered. The one limit is a courtesy to the sites being fetched: 20 checks a minute per visitor.

Questions

Frequently Asked Questions About the AI Crawler Checker

Which crawlers to allow, what blocking one really does, and where robots.txt stops.

Which AI crawlers should a regulated firm allow or block?

Split them by job. Allow the search crawlers and user fetchers, OAI-SearchBot, ChatGPT-User, PerplexityBot, Perplexity-User, Claude-SearchBot and Claude-User, because they are how an assistant finds and cites a page when someone asks. Make a deliberate call on the training crawlers, GPTBot, ClaudeBot and the Google-Extended token: allowing them suits educational content you want widely known, and blocking them suits writing that is itself the product or is covered by a client agreement. If you do block AI crawlers, block those names specifically rather than with a Disallow: / under *, which also shuts out Google Search. This is a content decision, not legal advice.

What is the difference between training bots and search or user-fetch bots?

A training bot collects pages that may be used to train future models; GPTBot and ClaudeBot are documented that way. A search bot builds the index an assistant searches when it answers, as OAI-SearchBot does for ChatGPT search and PerplexityBot does for Perplexity. A user-fetch bot, such as ChatGPT-User or Perplexity-User, requests one page because a person asked a question that needs it. OpenAI says robots.txt rules may not apply to ChatGPT-User, and Perplexity says Perplexity-User generally ignores them, because a user requested the fetch.

Does blocking GPTBot remove me from ChatGPT answers?

No. OpenAI documents GPTBot as its training crawler and says disallowing it does not affect ChatGPT search, which uses OAI-SearchBot; live fetches for a user come from ChatGPT-User. So a GPTBot block withholds your pages from training and leaves you eligible to be found and cited. Blocking OAI-SearchBot is what keeps you out of ChatGPT search answers, which OpenAI says may still show your site as a navigational link.

Is robots.txt enough, or can my firewall or CDN still block AI bots?

robots.txt is only the first gate. A CDN, web application firewall or security plugin can refuse a crawler by name or by IP before it ever reaches a page, and some offer AI-bot blocking as a setting, so the file can say allowed while the network says no. A page can also carry a noindex tag or an X-Robots-Tag header, or put its text behind JavaScript. This checker reads the file and cannot see any of that, so when the card says allowed and an assistant still cannot see you, the CDN and firewall settings are the next place to look.

Does Google-Extended affect Google rankings?

No. Google describes Google-Extended as a product token that controls whether content it crawls may be used to train future Gemini models, and says it "does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search". It is not a separate crawler; Googlebot does the fetching. Blocking it therefore has no effect on rankings, and the only robots.txt line that takes you out of Google Search is one that reaches Googlebot.

How often do AI crawler names change?

Often enough to re-read your file once in a while. Operators add crawlers as their products change: Anthropic now documents Claude-User and Claude-SearchBot alongside ClaudeBot, and anthropic-ai survives mainly as an older token in existing files. This checker tests 14 AI names, the same list Verand's site audit watches, including Claude-User and Claude-SearchBot, plus Googlebot for contrast. A misspelled or outdated name in your file blocks or allows nothing, so check each name against the operator's own documentation.

After the check

Open the right doors. Then give them something worth citing.

Access is the first gate, not the finish. Verand writes articles from your own expertise and credentials, so both Google and the AI assistants have something of yours to name. It then tracks where you rank and where ChatGPT, Gemini, Google AI Overviews, Google AI Mode, Perplexity and Claude mention you, and gates every draft so a claim your regulator would not allow never publishes.

Verand

Content built to rank in Google and get cited by ChatGPTPerplexityGeminiClaude, with every claim checked before it goes live.

support@verand.ai

© 2026 Verand. All rights reserved. TermsPrivacyAI policyAccessibilitySecurity
Not legal advice. Compliance packs are researched from the regulators' own text and tested by Verand, not reviewed by a licensed attorney.