Paste one URL. The checker reads that site's robots.txt for the exact path and tells you whether GPTBot, ClaudeBot, PerplexityBot and 12 other agents may fetch it, which rule decided each one, and what the page serves to a crawler that does not run JavaScript.
Two fetches, 15 agents. First /robots.txt at the root of the host, as VerandBot/1.0, redirects followed, 10 second timeout; a non-200 answer or an HTML body counts as no file. Then the page itself, once, 12 second timeout. The verdict covers the 14 AI agents Verand's site audit names, plus Googlebot for contrast.
Each agent takes the group that names it (the longest matching name wins), otherwise the * group. In that group the longest rule matching this path wins, and Allow wins a tie. * and $ in patterns are honoured. No file, or no group for the agent, means allowed.
robots.txt allowed does not prove the bot can reach the page. No request is sent as each bot, so a CDN, firewall or bot rule that turns one away does not show here. X-Robots-Tag headers are not read and JavaScript is not run.
One path, one file, one verdict per agent, then the page as it is served. The card is the tool in motion on an example investor page, looped, and each step lights up while the card is doing it.
200 · 1,180 words of main content · “Fund II closes to new investors on…”
The host tells the tool where the rules live: /robots.txt at its root, fetched as VerandBot with redirects followed. The path, query string included, is what every rule is tested against. A 404 or an HTML page at that address counts as no file.
An agent obeys the group that names it and ignores the rest. OAI-SearchBot has its own group here, so the * rules never reach it. GPTBot, ClaudeBot and PerplexityBot are not named, so they fall back to *.
Inside the chosen group, every Allow and Disallow pattern is tested against this one path. The longest match wins and Allow wins a tie. Disallow: /investors/ matches /investors/fund-ii/, so three agents are blocked from this page and only this part of the site.
Then the page itself is fetched once and read without running JavaScript: its status, whether it carries a noindex tag, and the main content with menus and footer stripped, which is the text a crawler that does not render would take away.
A site's robots.txt can be open at the root and closed on the page that matters, or the other way round. Here is how one URL gets its verdict, which bots ask, what they take away, and why a blocked page is sometimes the right answer.
Most AI crawler tools answer a site-level question: does robots.txt shut a bot out of the whole domain? That is the right first check, and the AI Crawler Checker runs it; the Robots.txt Checker reads the file itself line by line. Neither tells you about a particular page. A file that allows everything at / can still disallow /resources/, /clients/ or a single campaign URL, and a bot that is welcome on the home page may be turned away from the article you most want quoted.
This checker answers the narrower question. You give it one URL; it splits off the path, reads the site's robots.txt, and judges that exact path for each agent. Then it fetches the page once, so you also see what a crawler would receive if it were allowed in. Per-URL checking is not new (TechnicalSEO and serp.tools offer it, and both also send live requests as each bot, which this tool does not). What this page adds is the rule that decided each verdict, and a reading of the result as a policy you can confirm or change.
The file is a list of groups. Each group opens with one or more User-agent lines and carries Allow and Disallow patterns. RFC 9309, the standard written down in 2022, and Google's open-source parser decide a single URL in two steps.
First, the group. An agent looks for a group that names it. If one exists, the agent follows that group only and ignores every other line in the file, including the * group. If none names it, it follows *. If there is no * group either, or no file at all, nothing restricts it. This checker matches names without regard to case and takes the longest name that fits, so a group for GPTBot is not confused with one for ChatGPT-User.
Second, the rule. Inside that group every pattern is compared with the path, query string included. A pattern matches from the start of the path; * stands for any run of characters and a trailing $ anchors the end. Of the patterns that match, the longest one decides. When an Allow and a Disallow of the same length both match, Allow wins. An empty Disallow: matches nothing.
# path: /insights/fund-ii-update/ User-agent: * Disallow: /insights/ # matches, 10 characters Allow: /insights/fund- # matches, 15 characters: this rule wins User-agent: GPTBot Disallow: / # GPTBot reads only this group: blocked
So in that file ClaudeBot and PerplexityBot may read the page, because the longer Allow beats the shorter Disallow in the * group, and GPTBot may not, because its own group replaces * entirely. The card on this page prints that reasoning for each agent: which group it took, and the line that settled it.
Fifteen agents are judged on every run: the 14 AI user-agent tokens Verand's site audit watches, and Googlebot so you can compare the AI verdicts with the one that governs Google Search. The descriptions below are how each operator documents its agent, read on 27 September 2026. Check the operator's page before you block one, because the roles are not interchangeable.
| Agent | Operator | What the operator says it does |
|---|---|---|
| GPTBot | OpenAI | Crawls content that may be used to train OpenAI's foundation models. Disallowing it signals the content should not be used for training. |
| OAI-SearchBot | OpenAI | Surfaces websites in ChatGPT's search features. OpenAI says sites that disallow it will not be shown in ChatGPT search answers. |
| ChatGPT-User | OpenAI | Visits a page when a ChatGPT user's request calls for it. Because a person started the fetch, OpenAI says robots.txt rules may not apply. |
| ClaudeBot | Anthropic | Gathers web content that may be used to train Anthropic's models. Anthropic says its bots honour robots.txt. |
| Claude-SearchBot | Anthropic | Indexes content to improve Claude's search results. Anthropic says blocking it may reduce a site's visibility in those results. |
| Claude-User | Anthropic | Visits a page when someone asks Claude a question that needs it. Unlike OpenAI and Perplexity for their user fetchers, Anthropic says this one honours robots.txt. |
| anthropic-ai | Anthropic | An older token that no longer appears in Anthropic's current documentation. Checked because many files still carry a block for it. |
| PerplexityBot | Perplexity | Surfaces and links websites in Perplexity's search results. |
| Perplexity-User | Perplexity | Visits a page to answer a user's question. Perplexity says it generally ignores robots.txt, since a person asked for the fetch. |
| Google-Extended | Not a crawler. A robots.txt token that controls whether content Googlebot already fetched may be used to train future Gemini models. | |
| CCCCBot | Common Crawl | Builds Common Crawl's open repository of web crawl data, which researchers and companies reuse, including as training data. Honours robots.txt. |
| BBytespider | ByteDance | ByteDance's crawler. ByteDance publishes little about what it collects for, so treat it as a bulk crawler. |
| AAmazonbot | Amazon | Used to improve Amazon's products and services, and Amazon says the content may be used to train its AI models. Honours robots.txt. |
| MMeta-ExternalAgent | Meta | Crawls for uses such as training foundation AI models or improving products by indexing content directly. |
| Googlebot | The crawler behind Google Search. Included for contrast: it is the only agent in this list whose block removes a page from Google. |
The table falls into three jobs, and the difference is what makes a per-page decision possible. Training crawlers (GPTBot, ClaudeBot and the Google-Extended control, plus the bulk crawlers CCBot, Bytespider, Amazonbot and Meta-ExternalAgent) collect text that may end up in a future model. Search crawlers (OAI-SearchBot, Claude-SearchBot, PerplexityBot) build the index an assistant searches when it answers a question and links its sources. User fetchers (ChatGPT-User, Claude-User, Perplexity-User) retrieve a specific page because a person asked about it.
Operators document these as independent. OpenAI says outright that a site can allow OAI-SearchBot to appear in search results while disallowing GPTBot. That means a firm can keep its pages out of training sets and still leave them open to the crawlers that decide which pages an assistant can search. It also means a block written in a hurry, after a news story about AI training, often lands on the wrong agent. Whether any assistant then cites a page depends on far more than access, and nothing in robots.txt or on this card predicts it.
Many AI crawler checkers treat every block as a mistake and hand you the lines to undo it. For a business that publishes to be found, a site-wide block usually is a mistake. A block on one path often is not. A regulated firm may want its insights, guides and disclosures open to AI search, and its investor portal, offering documents, client pages and form confirmations closed to every crawler. A thin page that exists to say "thanks, check your inbox" has nothing to offer an answer engine, and a page meant only for accredited investors was never written for one.
That is why the card reads the result as a policy question: is this the access you intended for this page? Willowdale Equity's robots.txt shows the pattern in practice. It opens with Allow: / for every agent, then disallows its build folders, design mockups and newsletter thank-you pages. Checked here, its articles come back allowed for all 15 agents, and its /thank-you/ page comes back blocked for all 15 by Disallow: /thank-you/. Both answers are the ones the file was written to give.
Two limits apply to any page you close this way. The file is public, so a Disallow line tells everyone the path exists; anything confidential needs a login, not a robots rule. And a blocked page can still appear in Google as a bare address if other pages link to it, because Google was told not to read it, not told to forget it. To keep a page out of the index, it needs a noindex tag that a crawler is allowed to fetch; the Noindex Checker reads that tag. A page that is both disallowed and noindexed, like that thank-you page, is making a request crawlers cannot read.
robots.txt is an instruction, not a gate. It says what an agent is asked to do; it does not decide what the server answers. Two things can separate an "allowed" verdict from a bot actually receiving the page.
The first is infrastructure in front of the site. A CDN's bot management, a web application firewall or a hosting provider's security rule can answer a named AI agent with a 403 or a challenge page while the robots.txt says welcome. Some CDNs now offer one-click AI-bot blocking that lives entirely outside the file. This checker now fetches the page once with each crawler's user agent, beside the plain-browser request, and shows the answer as "reaches the page" next to each robots.txt verdict. That catches a block keyed on the user agent, which is how Cloudflare's AI bot policy works. A block keyed on the crawler's network addresses, which some hosts use for verified bots, is invisible to any test that is not the crawler itself; your CDN's firewall log is the definitive answer there.
The second runs the other way. OpenAI and Perplexity both say their user-triggered fetchers may not follow robots.txt, because a person asked for the page. So a Disallow line does not reliably keep ChatGPT-User or Perplexity-User out either. If a page must not be fetched by anyone, only authentication does that.
Access is half the question; the other half is what an agent takes away once it is in. The second half of the card shows the page the way a crawler that does not run JavaScript receives it: the HTML as served, with scripts and styles removed, and the main content separated from menus, header, footer and sidebars by the same rule Verand's site audit uses on every page it crawls.
That view matters because the AI crawlers have not behaved like browsers. Vercel's analysis of crawler traffic, published in December 2024, found that the major AI crawlers, OpenAI's and Anthropic's among them, fetched JavaScript files but did not execute them. A page that fills in its text after load can look complete in your browser and arrive nearly empty. A low count is not always a rendering problem, though. Verand's own help centre home page, checked here, returns 22 words of main content: a one-line intro and the names of its sections. It is served whole; it is simply a hub whose articles carry the text. The test is the comparison: if the main-content count on the card is far below what you see on screen, the text is arriving by JavaScript, and moving it into the served HTML is the fix.
A few limits of this reading. Whole-page word and byte figures come from the first 100,000 characters of HTML, so on heavier pages the card withholds them rather than print a short number; the main-content figures read the full page and stand. The text excerpt is capped at 20,000 characters in transit. And the noindex line reads the robots meta tag only, not an X-Robots-Tag response header.
robots.txt tells crawlers what they may request. llms.txt is a proposed file that lists a site's most useful pages for AI tools in plain Markdown. It grants or refuses nothing, and no major AI operator has confirmed that its systems read it. It is best treated as a curated index of the pages a firm has reviewed and wants to stand behind. The llms.txt Generator writes one from your own pages; this checker decides nothing from it.
Disallow: /resources with no trailing slash also blocks /resources-guide/ and /resources.html, because patterns match from the start of the path.Allow: /insights/ does let it read your insights, and also lifts every other * rule for it. Repeat the restrictions you still want inside its group.Disallow: /*? blocks every URL that carries a parameter, tracking links included./ says nothing about /clients/. Check the pages that carry your expertise, one at a time.Six things that are true of this tool, each one backed by a line in the code that runs it.
The request carries one URL and nothing else. There is no account, no session and no database behind the tool, so there is nothing for us to keep about you.
Every verdict comes back with the group the agent followed and the exact line that settled it, so a surprise block points straight at the line to change.
The file is judged by a fixed parser following RFC 9309's matching rules. No model reads it, so the same file and the same path give the same verdict every time.
The agent list and the main-content extraction are the ones Verand's site audit runs on every customer page, called directly. The page text you see is the text the product works from.
No live request per bot, no response headers, no JavaScript, no whole-page figures past the reading limit. The card says so beside the result rather than in a footnote.
Each run is two fetches and a parse, with paid page renders switched off, so it costs nothing and is never metered. The one limit is a courtesy to the sites being fetched: 20 checks a minute per visitor.
Which bots, which rules, and what an allowed verdict does and does not promise.
A GPTBot block does not close ChatGPT's search to the page. OpenAI documents GPTBot as its training crawler and OAI-SearchBot as the one that surfaces sites in ChatGPT search, and says the two settings are independent: a site can allow OAI-SearchBot while disallowing GPTBot. Whether ChatGPT then cites a page depends on much more than access, and neither robots.txt nor this checker can predict it. The card shows both agents separately so you can see which one your file reaches.
Training crawlers, such as GPTBot and ClaudeBot, collect text that may be used to train future models. Search crawlers, such as OAI-SearchBot and PerplexityBot, build the index an assistant searches when it answers a question and links sources. A third group, user fetchers such as ChatGPT-User and Perplexity-User, retrieve a page because a person asked about it, and their operators say robots.txt may not apply to them. Operators document the roles separately, so you can block one job without blocking the others.
Because robots.txt rules are path patterns. A file can allow everything at / and still disallow a folder such as /clients/ or a single URL, and the longest matching rule decides for each page. The card names the line that blocked the page. That is often deliberate: a firm that wants to block AI crawlers from investor-only pages, offering documents or form confirmations while keeping its articles open does it exactly this way. If the block is not deliberate, remove or narrow that line, and re-check the page.
No. robots.txt allowed does not prove the bot can reach the page. The file is a request to well-behaved crawlers; a CDN's bot management, a firewall or a security plugin can still answer a named AI agent with a block or a challenge. This checker reads the file and fetches the page once as VerandBot, and sends no request as each bot, so it cannot see a block of that kind. Tools that send requests with each bot's user agent, or your CDN's firewall log, can.
It follows the * group. An agent with no group of its own obeys the rules written for every crawler, so a Disallow under User-agent: * applies to it too. If there is no * group, or no robots.txt at all, nothing restricts it and the page is allowed. The card shows which group each agent fell into: its own name, *, or none.
No. Google documents Google-Extended as a robots.txt product token with no user agent string of its own: crawling is done by Google's existing crawlers, and the token controls whether content they fetched may be used to train future Gemini models. Google also says it does not affect inclusion or ranking in Google Search. So a Google-Extended block leaves Google Search alone, and a Googlebot block is the one that removes a page from Google.
Access only decides who may read the page. Verand writes articles from your own expertise and credentials, so what the crawlers find is yours. It then tracks where you rank on Google and where ChatGPT, Gemini, Google AI Overviews, Google AI Mode, Perplexity and Claude name you, and gates every draft so a claim your regulator would not allow never publishes.
Content built to rank in
Google and get cited by
ChatGPT
Perplexity
Gemini
Claude, with every claim checked before it goes live.