See whether ChatGPT, Claude, Perplexity and Gemini are allowed to read your site. One line in robots.txt can quietly shut them out of every page, so this free check tests 14 AI crawlers by name and puts the verdict and the fix right here.
2 checks. Is there a robots.txt, and does it block any of 14 AI crawlers site-wide. One GET to /robots.txt as VerandBot/1.0, redirects followed, 10 second timeout. A non-200 answer or an HTML body (the soft 404 many hosts serve) counts as no file.
The file is read in User-agent groups. A crawler is blocked when its own group, or the * group, contains Disallow: /. Blocked means the whole site; that is the only rule we score, because it is the one that stops an engine reading anything at all.
Path-level Disallows, Allow lines, wildcards and Crawl-delay are not evaluated. Nor are X-Robots-Tag headers or meta robots tags. If a bot wall blocks our fetch we report no file rather than buy a proxy render. Whether a crawler honours the file is the operator's promise, not ours to verify.
One fetch, one parser, one question per crawler. The card is the tool in motion on an example file, looped, and each step lights up while the card is doing it.
User-agent: GPTBot Allow: /
The tool asks for /robots.txt at the root of the host you typed, identifying itself as VerandBot, following redirects, giving up after ten seconds. A 404, a timeout or an HTML page served at that path all count as "no file".
Comments are stripped, then each User-agent line opens a group and every Disallow line beneath it is attached to that agent. Consecutive User-agent lines share one rule block, the way the protocol says they should.
For each of 14 named crawlers, and for *: is there a Disallow: / in its group? A yes is a site-wide block and a fail. Anything narrower is left alone on purpose, because a disallowed folder is usually deliberate.
The snippet on the card lists ten agents, the assistants' own crawlers and fetchers. A named group replaces * for that agent, so each group repeats your * rules and what you block stays blocked. A crawler you already name is left as written; paste only the ones you want.
The file is forty years younger than the web's oldest habits and still the first thing every crawler reads. Here is what it controls, what it does not, and the mistakes that cost sites their place in ChatGPT, Perplexity and Gemini.
A robots.txt is a plain-text file served at the root of a site, https://yourdomain.com/robots.txt, and nowhere else. It is the Robots Exclusion Protocol, drafted informally in 1994 and written down as an IETF standard, RFC 9309, in 2022. A crawler that respects the protocol fetches this file before it fetches anything else on the host and reads it as a set of instructions: which parts of the site that crawler may request, and which it should leave alone.
The file is built from groups. A group starts with one or more User-agent lines naming the crawler the group speaks to, and continues with Disallow and Allow lines listing path prefixes. A Sitemap line can sit anywhere and points crawlers at your XML sitemap. Comments begin with #. That is the whole grammar.
# every crawler not named below User-agent: * Disallow: /admin/ Disallow: /cart # one crawler, its own rules User-agent: GPTBot Disallow: / Sitemap: https://yourdomain.com/sitemap.xml
Two rules of the protocol decide almost everything. First, a crawler obeys the group that names it most specifically and ignores the rest: if there is a GPTBot group, GPTBot never reads the * group. Second, a crawler with no matching group at all has no restrictions. So a site with no robots.txt is fully open, and a robots.txt with an empty Disallow: is also fully open. Restriction only ever comes from an explicit line.
The file exists to manage crawl load and to keep crawlers out of pages that are useless to them: search results pages, faceted filters, staging paths, cart and checkout flows, endless calendar archives. It was never a security mechanism. Every crawler that honours it does so voluntarily, and the file is public, so a Disallow: /private-report/ line is a signpost, not a lock. Anything that must stay private needs authentication.
It is also not the way to keep a page out of Google's index. A URL blocked in robots.txt can still be indexed if other pages link to it; Google simply shows the address with no snippet, because it was told not to read the page. The directive that removes a page from the index is a noindex meta tag or an X-Robots-Tag header, and both only work if the crawler is allowed to fetch the page and see them. Google stopped honouring noindex written inside robots.txt in 2019. Blocking a page and asking for it to be de-indexed at the same time is the most common way to get neither.
AI assistants reach your site through named crawlers, and each operator publishes the name its crawler sends in the User-agent header. Those names are the whole point of this checker. A group like the GPTBot one above, two lines long, means OpenAI's training crawler requests nothing from your site. A Disallow: / under User-agent: * means every crawler that is not given its own group, which is usually all of them, stops at the door. On a site that publishes expertise so that people find it, that second case shuts out every AI crawler that honours the file, and the file does not warn you.
It happens quietly for three reasons. Blanket blocks are often installed by a hosting control panel, a security plugin or a developer protecting a staging copy, then copied to production. Named blocks are added when a story about AI training makes the rounds, without separating the crawlers that train models from the ones that fetch a page to answer a live question. And nothing in a normal analytics stack reports a crawler that never arrived. The site keeps ranking in Google, which uses a different crawler, so the block is invisible right up until someone asks an assistant a question you should have answered.
These are the names Verand's site audit watches on every customer site, and the checker tests each one against your file. The role each plays is how its operator documents it; check the operator's page before you decide to block one, because the roles are not interchangeable.
| Crawler | Operator | What it does |
|---|---|---|
| GPTBot | OpenAI | Crawls the open web for model training. Blocking it does not touch ChatGPT's live browsing or search. |
| ChatGPT-User | OpenAI | Fetches a page when a ChatGPT user's question calls for it. Because a person started the fetch, OpenAI says robots.txt rules may not apply to it. |
| OAI-SearchBot | OpenAI | Builds the index behind ChatGPT search. Block it and your pages are not candidates for search results inside ChatGPT. |
| ClaudeBot | Anthropic | Anthropic's training crawler: it collects web content that may be used to train Anthropic's models. Blocking it keeps future pages out of training data. |
| Claude-User | Anthropic | Fetches a page when someone asks Claude a question that needs it. Anthropic says it honours robots.txt, so blocking it stops Claude retrieving your pages for users. |
| Claude-SearchBot | Anthropic | Indexes content to improve Claude's search results. Anthropic says blocking it may reduce your visibility in them. |
| anthropic-ai | Anthropic | An older Anthropic token still present in many robots.txt files. Kept in the list so a legacy block is not missed. |
| PerplexityBot | Perplexity | Builds Perplexity's search index. Perplexity recommends allowing it if you want your site to appear in its search results. |
| Perplexity-User | Perplexity | Fetches a page in response to a user's specific question. Perplexity says it generally ignores robots.txt, because a user asked for the fetch. |
| Google-Extended | A control token, not a separate crawler: it governs whether Googlebot's crawl may be used for Gemini training and grounding. It does not affect Google Search. | |
| CCCBot | Common Crawl | Builds Common Crawl's open repository of web crawl data, which researchers and companies reuse, including as training data. |
| BBytespider | ByteDance | ByteDance's crawler, associated with its models and products. |
| AAmazonbot | Amazon | Used to improve Amazon's products and services; Amazon says the content may be used to train its AI models. |
| MMeta-ExternalAgent | Meta | Crawls for uses such as training foundation AI models or improving products by indexing content directly. |
The fix this page hands you lists the ten of those that belong to OpenAI, Anthropic, Perplexity and Google, whose assistants Verand tracks. The other four gather pages in bulk for datasets, model training and product indexes; they are reported so you know they are blocked, and whether to allow them is a policy choice about training data, and a fair one to make either way.
The confusion that produces most bad robots.txt files is treating Google and the AI assistants as one audience. They are not. Googlebot is the crawler behind Google Search, and blocking it is the only way to fall out of Google. Google-Extended is a separate token that controls whether the pages Googlebot already fetched may be used for Gemini; Google documents that it has no effect on ranking or on inclusion in Search, and that AI Overviews and AI Mode are part of Search and draw on the same index. So a Google-Extended block does not remove you from Google's AI answers, and a GPTBot block does nothing to your Google rankings.
The practical consequence: a site can be perfectly visible in Google while being absent from ChatGPT and Perplexity, and the owner will see healthy Search Console numbers the whole time. The robots.txt is where that split is decided, one User-agent group at a time, which is why this checker reports by name rather than with a single pass or fail.
User-agent lines above it; only a User-agent line that comes after rules starts a new group. Some older parsers did treat a blank line as a break, so keep each group's lines together and never rely on spacing to decide which crawler a rule applies to. Rules placed before any User-agent line belong to no group and do nothing.GPT-Bot is not GPTBot. A misspelled group blocks nothing and gives false comfort.For most sites that publish expertise: a * group that keeps crawlers out of the paths no reader should land on, a Sitemap line, and no site-wide Disallow for any crawler you want citing you. If you have decided to withhold training data from a particular operator, block its training crawler by name and leave its fetcher and search crawler open, so the decision is precise rather than a blanket outage. Then re-check after every deploy, because the file is the kind of thing that gets overwritten.
Six things that are true of this tool, each one backed by a line in the code that runs it.
The request carries a domain and nothing else. There is no account, no session and no database behind the tool, so there is nothing for us to keep about you.
Every response includes the allow lines for ten AI crawlers, the assistants' own. Nothing is held back for a report, a call or an upgrade.
The file is parsed by a fixed grammar and matched against a fixed list of 14 names. No model reads it, so the same file gives the same answer every time.
This is the site-level check Verand's deep crawl runs on every customer site on the 1st and 15th of each month, called directly. Not a lighter demo version.
Only a site-wide Disallow is scored. Path rules, Allow lines, headers and meta tags are out of scope, and the card says so beside the result rather than in a footnote.
Each run is two small fetches, so it costs nothing and is never metered. The one limit is a courtesy to the sites being fetched: 20 checks a minute per visitor.
What it tests, what a missing file means, and what it never does to your site.
Two things. Whether a robots.txt exists at the root of the domain and is served as plain text, and whether that file blocks any of 14 named AI crawlers from the whole site with a Disallow: / line, either in the crawler's own group or in the * group. It reads the file in User-agent groups the way a crawler does, with consecutive User-agent lines sharing one block and comments stripped. It does not evaluate path-level rules, Allow lines, wildcards or Crawl-delay.
It is reported as a fail at the lowest severity, and the card says why: a missing file allows all crawling, so nothing is blocked, but nothing is controlled either. You have no Sitemap line, no way to keep crawlers out of search-result or checkout pages, and no place to record a deliberate decision about any AI crawler. The fix is a short file, and the allow lines on the card are a reasonable start for one.
No. Google Search uses Googlebot, and a GPTBot group speaks only to OpenAI's training crawler. The reverse is also true: blocking Google-Extended does not remove you from Google Search, AI Overviews or AI Mode, because Google documents that token as a control on Gemini training and grounding, not on Search. The only line that removes you from Google is a Disallow that reaches Googlebot, which a site-wide Disallow: / under * does.
No. It makes two GET requests from our server, one for robots.txt and one for llms.txt, and writes nothing. The fix is text on the page for you to paste; the checker has no access to your site and asks for none. It also refuses to fetch private or internal addresses, so it cannot be pointed at anything behind your firewall.
Because only a site-wide block is scored. Disallow: /admin/ or Disallow: /search under a crawler's name keeps it out of those paths and nothing else, and the crawler can still read and cite the rest of the site, so the checker reports it as allowed. The one case to know about: a * group that says Disallow: / is reported as blocking every crawler, even when a named group beneath it says Allow: /, because Allow lines are not evaluated. If you have added the fix and the card still shows a block, the * group is the line to change.
Free, with no account and no daily allowance. A run is two small fetches and a parse, so it costs us nothing and there is no meter. The one limit is 20 checks a minute per visitor, which is a courtesy to the sites being fetched rather than a quota. Run it after every deploy; the file gets overwritten more often than anyone expects.
An open robots.txt is the first line, not the finish. Verand writes articles from your own expertise and credentials, so both Google and the AI assistants have something of yours to name. It then tracks where you rank and where ChatGPT, Gemini, Perplexity, Claude and Google's AI answers mention you, and gates every draft so a claim your regulator would not allow never publishes.
Content built to rank in
Google and get cited by
ChatGPT
Perplexity
Gemini
Claude, with every claim checked before it goes live.