Decide, crawler by crawler, which AI bots may read your site: the ones that train models, the ones that build the search indexes behind ChatGPT, Claude and Perplexity answers, and the ones that fetch a page because someone asked. Start from your live file and every group it adds repeats your existing rules, so nothing private opens by accident.
One group per crawler you decide on. Allow repeats your User-agent: * rules, and its first Crawl-delay, or writes Allow: / when there are none. Block writes Disallow: /. No rule writes nothing, so that crawler follows *. From your live file, the file comes back whole with the new groups appended; from scratch, you get a complete file.
14 names, from the list Verand's own site audit watches, grouped by what each operator's documentation says the crawler is for. Where the operator says nothing, it sits under Other rather than being guessed at. A crawler your file already names is left exactly as written. Googlebot, Bingbot and other search engines are not in the list.
Make any crawler obey. robots.txt is a request: OpenAI says it "may not apply" to ChatGPT-User, and Perplexity says Perplexity-User "generally ignores" it. A CDN or firewall can still block a crawler you allowed. Blocking a crawler stops it fetching pages, not an engine citing you from somewhere else. Review the file before you upload it.
Load the file you have, pick what each crawler may do, and get groups that keep your existing rules. The card is the tool in motion on an example domain, looped, and each step lights up while the card is doing it.
User-agent: GPTBot Disallow: / User-agent: OAI-SearchBot Disallow: /wp-admin/ Disallow: /client-portal/
Type your domain and the tool reads /robots.txt once, as VerandBot. Skip it and you build from scratch, with nothing sent anywhere. Either way the file below is assembled in your browser.
Every Allow and Disallow line under User-agent: * is collected in file order. These are the lines a crawler loses the moment it gets a group of its own.
Pick a preset or set each crawler yourself: allow, block, or no rule. Training, search and user-fetch crawlers are listed apart, each with what its operator says about robots.txt beside it.
Each Allow group repeats your general rules, so /client-portal/ stays closed to a crawler you let in. Block groups are Disallow: /. Copy the file or download it, then check it's live.
A robots.txt for AI is not one line that says yes or no to "AI". It is a set of decisions about separate crawlers that do separate jobs, written so the rest of your file keeps working. Here is what each decision costs and where the file stops.
robots.txt is a plain text file at the root of a host that tells crawlers which paths they may fetch. It is organised in groups: a User-agent line naming a crawler, then the Allow and Disallow rules that apply to it. User-agent: * is the group for everyone not named elsewhere. AI crawlers read the same file as search engines, so a robots.txt for AI is the same file with groups for the AI crawlers you want to treat differently.
Two things follow from that. The file is a request, honoured by crawlers that choose to honour it, not a lock; anything that must stay private needs a login. And it works per host, so www.yourdomain.com, the bare domain and help.yourdomain.com each serve their own file. If you want to see what a host's current file already says to each AI crawler, the AI Crawler Checker gives a verdict per bot; this page writes the policy.
Each major AI operator now runs more than one crawler, and they do different work. A training crawler collects pages that may be used to train future models. A search crawler builds the index an assistant searches when it composes an answer. A user fetcher requests one page in the moment, because a person asked a question that needs it. Blocking one leaves the others exactly where they were. These are the 14 names the builder covers, grouped the way their operators describe them:
| Crawler | Operator | Caveat the operator states |
|---|---|---|
| Training | ||
| GPTBot | OpenAI | No caveat stated |
| ClaudeBot | Anthropic | No caveat stated |
| Google-Extended | A robots.txt control token, not a separate crawler; Google: it "does not impact a site's inclusion in Google Search nor is it used as a ranking signal". | |
| Amazonbot | Amazon | Amazon: also used "to improve our products and services". |
| Meta-ExternalAgent | Meta | Meta: also used for "improving products by indexing content directly". |
| Search | ||
| OAI-SearchBot | OpenAI | No caveat stated |
| Claude-SearchBot | Anthropic | No caveat stated |
| PerplexityBot | Perplexity | No caveat stated |
| User fetch | ||
| ChatGPT-User | OpenAI | OpenAI: "Because these actions are initiated by a user, robots.txt rules may not apply." |
| Claude-User | Anthropic | No caveat stated |
| Perplexity-User | Perplexity | Perplexity: "Since a user requested the fetch, this fetcher generally ignores robots.txt rules." |
| Other | ||
| anthropic-ai | Anthropic (legacy token) | Not listed on Anthropic's current crawler page; kept because older robots.txt files still name it. |
| CCBot | Common Crawl | Builds Common Crawl's open web archive, which anyone can download; Common Crawl's own pages name no AI-training use. |
| Bytespider | ByteDance | ByteDance publishes no documentation for this crawler, so neither its purpose nor its robots.txt handling is confirmed by the vendor. |
The grouping is read from each operator's own documentation. Where an operator publishes nothing, as ByteDance does not for Bytespider, the crawler sits under Other and the note says so, rather than carrying a guess. The same list drives Verand's site audit, so a name added there appears here on the next build of this page.
The most common request is to keep writing out of training sets while staying findable when a client asks an assistant a question. The preset "Allow AI search and user fetches, block training" writes exactly that: Disallow: / for GPTBot, ClaudeBot, Google-Extended, Amazonbot and Meta-ExternalAgent, and allow groups for OAI-SearchBot, Claude-SearchBot, PerplexityBot, ChatGPT-User, Claude-User and Perplexity-User. The Other crawlers (the legacy anthropic-ai token, CCBot and Bytespider) get no rule, so they follow your * group until you decide.
Why this split holds up: OpenAI says disallowing GPTBot "indicates a site's content should not be used in training generative AI foundation models", while sites that opt out of OAI-SearchBot "will not be shown in ChatGPT search answers". Perplexity says PerplexityBot "is not used to crawl content for AI foundation models". Anthropic describes ClaudeBot as collecting content that "could potentially contribute" to training, and Claude-SearchBot as improving "search result quality". Blocking the training names and allowing the search names is therefore two separate decisions, and each operator documents them separately. Other tools offer the same preset; the parts that decide whether your file still works are below.
Under RFC 9309, the robots.txt standard, a crawler obeys only the group that names it. If a group names GPTBot, GPTBot ignores User-agent: * entirely, including every Disallow in it. That is well known on the block side. It is rarely mentioned on the allow side, where it does the damage: adding a plain allow group opens every path your general rules closed.
# before User-agent: * Disallow: /wp-admin/ Disallow: /client-portal/ # a common "allow GPTBot" snippet, appended User-agent: GPTBot Allow: /
After that append, GPTBot may fetch /client-portal/ and /wp-admin/, because its own group says Allow: / and nothing else. The fix is to repeat the general rules inside the named group, which is what this builder does for every group it adds:
User-agent: GPTBot Disallow: /wp-admin/ Disallow: /client-portal/
Willowdale Equity's live file shows why this matters on a real site. Its User-agent: * group carries an Allow: / and 13 Disallow lines, for build folders, design mockups and thank-you pages. A bare allow group for an AI crawler would open all 13 to it. Loaded here, every allow group repeats all 14 lines. The robots.txt fix inside Verand's product writes its groups the same way.
If your * group has a Crawl-delay, it is repeated too, since a named group loses that as well. And if your * group says Disallow: / for the whole site, the builder repeats that faithfully and warns you, because an allow group that repeats a site-wide block allows nothing, and the real problem is the line under *.
When you start from your live file and it already has a group for, say, ClaudeBot, the builder does not add a second one. It shows that crawler as named in your file, with its current rule, and leaves it alone. The reason is that RFC 9309 merges two groups that name the same crawler, so a second group could quietly reverse the first, and a named group is usually a decision somebody made on purpose: counsel, a previous agency, or a CDN setting. To change it, edit that group in your file. The output says which crawlers were left and why, as a comment.
A generator that lists every crawler as if a Disallow works on all of them is telling you something the operators do not. The notes beside each crawler in the builder come from the operators themselves:
Generic presets assume a publisher. A firm that publishes expertise so prospects can find it has a different question: which machines can read our disclosures, and did someone block the assistants our prospects use to shortlist firms? A reasonable starting point by firm type, as a content decision and not legal advice:
Whatever you choose, write it down and re-check it after every site deploy, because robots.txt is the file that gets overwritten by a theme, a plugin or a staging copy.
Changes are not instant. OpenAI says "it can take ~24 hours from a site's robots.txt update for our systems to adjust" for search results, and Perplexity says it "may take up to 24 hours" for its systems to reflect a change. Then confirm the file is live and says what you meant with the AI Crawler Checker, check one important URL with the AI Crawler Access Checker, and read the whole file line by line in the Robots.txt Checker. For the non-AI parts of the file, search engines and sitemaps, the Robots.txt Generator covers them.
anthropic-ai survives mainly in older files.Disallow: / under * to stop AI. It stops Googlebot too, which takes the site out of Google Search.Six things that are true of this tool, each one backed by the code that runs it.
Every allow group repeats your User-agent: * rules and its crawl delay, so letting a crawler in never opens a path you had closed. The same rule the product's own robots.txt fix follows.
A crawler your file already names is shown with its current rule and not touched. No second group is added to merge with it and reverse it.
ChatGPT-User, Perplexity-User and the Google-Extended token carry their operators' own words beside them, so a block that is only a request reads as one.
The 14 names, their purposes and notes are the registry Verand's site audit reads, copied into this page when it is built. No model classifies anything, and the same choices always write the same file.
The builder sends nothing. The only request is the optional read of your live robots.txt, which carries a domain and nothing else. No account, no session.
Building a file costs nothing and is never metered. Loading a live file is limited to 20 reads a minute per visitor, as a courtesy to the sites being read.
Training versus answers, the crawlers robots.txt cannot bind, and what happens to the rules you already have.
Yes, because the operators run separate crawlers for each job. Block the training crawlers (GPTBot, ClaudeBot, Google-Extended, Amazonbot, Meta-ExternalAgent) and allow the search crawlers (OAI-SearchBot, Claude-SearchBot, PerplexityBot) and the user fetchers. OpenAI says disallowing GPTBot signals your content should not be used for training, and that only sites opting out of OAI-SearchBot are left out of ChatGPT search answers. The preset "Allow AI search and user fetches, block training" writes this split. Allowing a crawler makes you eligible to be read; it does not guarantee a citation.
Don't expect it to. Google-Extended is a control token, not a crawler: Googlebot fetches the page, and the token governs whether that content may be used for Gemini training and for grounding in Gemini Apps and Vertex AI. Google says it does not affect inclusion in Google Search and is not a ranking signal, and its description does not name AI Overviews or AI Mode, which are Search features. The only robots.txt line that takes a page out of Google Search is one that reaches Googlebot, and that removes it from ordinary results too.
Because a person asked for the page. OpenAI says of ChatGPT-User: "Because these actions are initiated by a user, robots.txt rules may not apply." Perplexity says Perplexity-User generally ignores robots.txt rules for the same reason. A block for either is a stated preference, not a control. Anthropic is the exception among the three: it says its bots, Claude-User included, honour robots.txt. If a page must not be fetched at all, it needs a login or a network rule, not a robots.txt line.
Not unless the new group repeats them. Under RFC 9309 a crawler follows only the group that names it and ignores User-agent: * entirely, so a plain "User-agent: GPTBot / Allow: /" opens every path your general rules closed, such as /wp-admin/ or a client portal. This builder repeats your User-agent: * rules, and its crawl delay, inside every allow group it writes, and it leaves alone any crawler your file already has a group for.
Most firms that publish to be found should allow the search crawlers and user fetchers, since those are how an assistant finds and quotes a page when a prospect asks, and it lets the assistant read your disclosures along with your claims. Training is a separate decision: allowing it suits educational writing you want widely known, and blocking it suits original writing that is itself the product. Anything that must stay private belongs behind a login, not in robots.txt. This is a content decision, not legal advice; your counsel decides what a client agreement or licence requires.
About a day, by their own documentation. OpenAI says it can take about 24 hours from a robots.txt update for its search systems to adjust, and Perplexity says it may take up to 24 hours for its systems to reflect a change. Anthropic's crawler page gives no figure. Once the file is uploaded, the AI Crawler Checker shows what it now says to each crawler straight away, even before the operators have re-read it.
Verand writes from your own experience, gets you found on Google and in AI answers, and tracks where you rank and which answers name you. Every draft passes a compliance gate before anyone can publish it: your industry's rules if you have them, a truth-in-advertising check if you don't.
Researched from the regulators' own text and tested by Verand. Not reviewed by a licensed attorney. Your counsel confirms applicability. Not legal advice. Example shown is illustrative.
The 2019 fund returned 14.2% net to investors, and this offering carries guaranteed returns of 12 to 15% over a five year hold.
Performance advertising can’t promise a return. The claim needs a basis and the required disclosures, and it still can’t be stated as a guarantee.
Willowdale Equity and four competitors · 14 days
Content built to rank in
Google and get cited by
ChatGPT
Perplexity
Gemini
Claude, with every claim checked before it goes live.