Find any site's XML sitemap the way a crawler looks for it: the Sitemap: lines in robots.txt first, then the standard addresses. See where it lives, whether robots.txt declares it, and how many URLs it lists, so you can decide whether that is the list you meant to publish.
2 checks and a count. Is there a sitemap, and does robots.txt declare it. The first sitemap that answers is read for what it lists: URLs in a urlset, child sitemaps in an index. Every request goes out as VerandBot/1.0, follows redirects and gives up after 10 seconds.
Up to three full-URL Sitemap: lines from robots.txt, then /sitemap.xml, /sitemap_index.xml, /wp-sitemap.xml and /sitemap-index.xml on the host you typed. An error status, an HTML page, or a file with no urlset or sitemapindex root is skipped and the next address is tried.
It does not validate the file: protocol problems are counted, not explained, and the Sitemap Validator does that job. Child sitemaps are listed, not opened, so an index's URL total is unknown here. Only the first ten entries are shown, and a file is read to 8 MB. It cannot tell which pages you have retired, or what Search Console has read.
robots.txt first, then the usual addresses, and it stops at the first real sitemap. The card is the tool in motion on an example site, looped, and each step lights up while the card is doing it.
Sitemap: https://example.com/sitemap_index.xml
The tool asks for /robots.txt on the host you typed and reads every Sitemap: line in it. Up to three that are full URLs go to the front of the queue, because a declared sitemap is the one the site owner meant crawlers to read.
/sitemap.xml, /sitemap_index.xml, /wp-sitemap.xml, /sitemap-index.xml, in that order. A 404, a timeout or an HTML page served at the address counts as no answer, and the next one is tried.
The first file with a urlset or sitemapindex root ends the search. Its entries are counted, URLs for a urlset and child sitemaps for an index, and the first ten are listed so you can see what it tells crawlers to read.
A sitemap found by probing works for crawlers that guess the same address. The card says which way it was found, and when robots.txt does not name it, the page writes the one line that does, with the address we found.
A sitemap is the one page on your site written for machines only. Here is what it does, where crawlers look for it, why robots.txt should name it, and why the list itself deserves a read before anyone else reads it.
An XML sitemap is a file that lists the URLs on a site you want search engines to know about. The format is the Sitemaps protocol at sitemaps.org, published jointly by Google, Yahoo and Microsoft in 2006, and it has barely changed since. Each entry is a <url> element holding a <loc>, the full address, and optionally a <lastmod> date for when the page last changed in a way that matters.
<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<url>
<loc>https://yourdomain.com/services/estate-planning/</loc>
<lastmod>2026-08-14</lastmod>
</url>
</urlset>
A sitemap is a discovery aid. Google describes it as a hint about which pages exist and which changed, not an instruction: listing a page does not make Google crawl it, and it does not make Google index it or rank it. What the file does well is shorten the path to pages your internal links reach slowly or not at all, such as a new article, a deep archive page, or a page that only a form or a search box leads to. Google also says it ignores the protocol's <priority> and <changefreq> fields, and reads <lastmod> only on sites where the dates prove consistently accurate.
There is no single required address, which is why a sitemap finder has to look in more than one place. The first place is the site's own robots.txt: a Sitemap: line there is the owner saying exactly where the file is, and it is the one lookup that never has to guess. After that come the addresses platforms use by default. This checker tries them in the order below, on the exact host you typed, and stops at the first that returns a real sitemap.
| Address | Who usually serves it |
|---|---|
Sitemap: lines in robots.txt | Any site whose owner declared the sitemap URL. It can be any path, which is why it is read first. |
/sitemap.xml | The conventional address. Shopify, Squarespace, Wix and most static site generators serve the sitemap here. |
/sitemap_index.xml | WordPress sites running Yoast SEO or Rank Math, which split the sitemap into an index of child files. |
/wp-sitemap.xml | WordPress core, which has generated its own sitemap at this address since version 5.5 when no plugin replaces it. |
/sitemap-index.xml | Some generators write the hyphenated name, the Astro sitemap integration among them. |
To find the sitemap URL by hand, open /robots.txt in a browser and look for the line; if it is not there, try the addresses above. For your own site, Search Console's Sitemaps report shows every sitemap submitted for the property and when Google last read each one. The host matters: www.yourdomain.com and yourdomain.com are separate hosts to a crawler, and each serves its own robots.txt.
The declaration is one line: the field name, a colon and the sitemap's full URL. It is not part of any User-agent group, so it can sit anywhere in the file and every crawler that reads robots.txt sees it, whatever rules apply to it. A site can list several, one per line, which is how a site with separate sitemaps for articles, locations and images points at all of them.
# access rules for crawlers User-agent: * Disallow: /client-portal/ # the declaration, outside any group Sitemap: https://yourdomain.com/sitemap_index.xml
The URL must be absolute, with the scheme and host written out. A relative path such as Sitemap: /sitemap.xml is not valid under the protocol, and this checker does not follow one. Declaring the file is not required by Google, but it is the only way to reach every crawler at once without an account with each search engine, which is why the check for it exists at all. For how the rest of the file controls access, the Robots.txt Checker reads it group by group.
A sitemap index is a sitemap of sitemaps: a <sitemapindex> root whose entries point at other sitemap files instead of pages. The protocol caps one sitemap file at 50,000 URLs and 50 MB uncompressed, so large sites have to split, and most CMSs split by content type anyway, one child for posts, one for pages, one for authors. When this checker finds an index, it reports how many child sitemaps it lists and shows the first ten addresses; it does not open them, so the number of pages behind an index is not known from this card. Checking each child's structure and limits is what the Sitemap Validator is for.
Beyond the robots.txt line, you can submit the sitemap URL directly: in Google Search Console under Sitemaps, and in Bing Webmaster Tools under Sitemaps. Submission is useful for your own reporting, because each tool then shows when it last read the file and how many URLs it found, but it is a second route to the same file rather than a replacement for the declaration. Google retired its old sitemap "ping" endpoint in 2023, so there is no longer an anonymous way to tell Google a sitemap changed; it rereads files it already knows on its own schedule.
The most common way to have a sitemap that does nothing is to have one that does not load. A declared address that returns 404 after a redesign. A sitemap URL that redirects to the homepage, so the crawler receives an HTML page where it expected XML. A security plugin or CDN rule that answers crawlers with a challenge page. A staging robots.txt that shipped to production still pointing at the staging host. This checker treats every one of those as no answer and moves on to the next address, and when robots.txt names a sitemap that did not load, the card says so by name rather than quietly reporting the one it found elsewhere.
Most sitemap tools treat a bigger URL count as a better result. For a firm in a regulated industry, the count is a different question. A sitemap is a public inventory of every page the firm is still asking search engines, and the crawlers behind AI answers, to read. Generated sitemaps list whatever the CMS still holds as published, which means they faithfully keep listing the pages nobody meant to keep: a withdrawn offering, a departed advisor's bio, last year's fee schedule, a disclosure that has since been superseded, a landing page for a campaign that ended.
That is why this card leads with the number and the first entries rather than a pass mark. If an adviser site you expect to hold forty pages lists three hundred, or the first entries include pages you retired, the sitemap is telling crawlers to keep reading something the firm no longer stands behind. Taking a page out of the sitemap does not remove it from Google; a retired page should return 404 or 410, or redirect to its replacement, and then drop out of the sitemap on its own. The sitemap is simply where the gap between what you publish and what you meant to publish shows up first.
/sitemap.xml find it, but nothing declares it, and a sitemap at any other address is invisible to them.Sitemap: /sitemap.xml is not a URL under the protocol. Write it out in full.www pointing at a sitemap full of apex URLs, or the reverse. Pick one host and use it everywhere.<lastmod> to today on each build teaches Google to ignore the field for the whole site.Six things that are true of this tool, each one backed by a line in the code that runs it.
The request carries a domain and nothing else. There is no account, no session and no database behind the tool, so there is nothing for us to keep about you.
The card names the address that answered, every address tried before it, and what robots.txt says. When the declaration is missing, the page writes the Sitemap: line from the URL we found, labelled as the page's own, since the tool returns no fix text for this check.
A fixed order of addresses, a fixed rule for what counts as an answer, and a fixed parse of the file. No model reads anything, so the same site gives the same answer until the site changes.
The missing-sitemap and not-in-robots.txt findings come from the same site-level function Verand's deep crawl runs on every customer site every two weeks. The probe that feeds it is this tool's own, written to look where crawlers look.
It finds and counts; it does not validate, open child sitemaps, or know which pages you retired. The card says so beside the result, and points at the tool that does the next job.
Each run is a handful of small fetches, so it costs nothing and is never metered. The one limit is a courtesy to the sites being fetched: 20 checks a minute per visitor.
Finding a sitemap, declaring it, and what the Search Console errors mean.
Start with the site's robots.txt: open yourdomain.com/robots.txt and look for a line beginning Sitemap:, which gives the exact URL. If there is none, try the addresses platforms use by default: /sitemap.xml, /sitemap_index.xml (Yoast and Rank Math on WordPress), /wp-sitemap.xml (WordPress core) and /sitemap-index.xml. This checker does both in that order and stops at the first address that returns a real sitemap. For your own site, Search Console's Sitemaps report also lists every sitemap submitted for the property.
No. Google accepts a sitemap submitted in Search Console, and crawlers that guess /sitemap.xml find a file at that address without help. The declaration is still worth the one line: it is the only route that reaches every crawler that reads robots.txt, without an account with each search engine, and it is the only way a sitemap at an unusual address gets found at all. This checker reports a missing declaration as a low-severity finding for that reason, and writes the line for you from the address it found.
It appears with the "Couldn't fetch" status in the Sitemaps report and means Google did not get a usable file the last time it tried. The usual causes are a URL that returns an error or a login page, a redirect to the homepage, a robots.txt rule that blocks Googlebot from the file, or a firewall or CDN challenge. It can also show for a while on a freshly submitted sitemap. Run the domain here: if the checker finds the file at the address you submitted, it loads for a normal server request, and the Sitemap Validator will then check its format. Our requests identify as VerandBot, not Googlebot, so a firewall that treats the two differently can give different answers.
A sitemap index is a file that lists other sitemap files instead of pages. The protocol caps a single sitemap at 50,000 URLs and 50 MB uncompressed, so very large sites must split, and many CMSs split by content type regardless of size. A small site does not need one; if your platform produces one, submit and declare the index and the children come with it. This checker reports how many child sitemaps an index lists and shows the first ten. It does not open them; the Sitemap Validator does.
Often not strictly. Google's own guidance says a site of about 500 pages or fewer, with every important page linked from somewhere else on the site, may not need one. It is still cheap to have, most platforms generate it without being asked, and it gives you one explicit list of the pages you intend to publish, which is useful to read in its own right. A new site with few inbound links benefits most, because its internal links are the only other way in.
Whenever a page is added, removed or meaningfully changed. A generated sitemap does this on its own; a hand-written one needs updating with each change. The lastmod date should move only when the content does, because Google uses it only on sites where the dates prove accurate. There is no need to resubmit after each change: Google retired its sitemap ping endpoint in 2023 and rereads sitemaps it already knows about on its own schedule.
Finding the file is the first step. Verand writes articles from your own expertise and credentials, so Google and the AI assistants have something of yours worth naming. It then tracks where you rank and where ChatGPT, Gemini, Perplexity, Claude and Google's AI answers mention you, and gates every draft so a claim your regulator would not allow never publishes.
Content built to rank in
Google and get cited by
ChatGPT
Perplexity
Gemini
Claude, with every claim checked before it goes live.