Enter a sitemap's address, or just your domain to check its /sitemap.xml. The validator fetches the file, checks it against the sitemaps.org protocol, names every error and says whether it makes the whole file invalid or only drops one entry, then counts listed URLs your own robots.txt tells Google not to crawl.
Up to 10 checks on one file. One GET for the URL you enter, as VerandBot/1.0, redirects followed, 10 second timeout. The root element and its sitemaps.org 0.9 namespace; entries present and under 50,000 URLs; every entry has a <loc> that is absolute and on the sitemap's own host; <lastmod> in W3C format; no duplicate <loc>. Each error is marked as one that invalidates the file or one that drops an entry.
For a sitemap index, the first five child sitemaps are fetched and each is summarised with its URL count and error count. For a URL set, the first 200 listed URLs are tested against the site's own robots.txt as Googlebot: the longest matching rule wins, and Allow beats Disallow on a tie.
The file is read by pattern matching, not a full XML parser, so a malformed tag or an unescaped & can pass. It reads the first 8 MB, so the 50 MB limit is not measured on larger files. Gzipped .xml.gz files are not unpacked. <changefreq>, <priority>, image and video tags, the UTF-8 encoding and noindex tags on the listed pages are not checked. A pass means the file meets these checks, not that Google has accepted it.
One fetch, one pass over the entries, one question per rule. The card is the tool in motion on an example file with two mistakes in it, looped, and each step lights up while the card is doing it.
<loc>https://example.com/about</loc> <lastmod>2026-09-12</lastmod>
The tool asks for the URL you entered, or for /sitemap.xml when you enter only a domain, identifying itself as VerandBot, following redirects, giving up after ten seconds. Anything but a 200 stops here with the status code, because there is nothing to validate.
A <urlset> root makes it a list of pages; a <sitemapindex> root makes it a list of sitemaps. An HTML page, an empty body or neither root ends the run with one file-level error.
Each entry is read for its <loc> and <lastmod>, and the errors are counted per rule. Then up to 200 of the listed URLs are matched against the site's robots.txt, the way Googlebot would read it.
Every error on the card says whether the protocol treats the whole file as invalid or whether the file is still read and only that entry is dropped. Where a fix is one line, the card prints it.
A sitemap is a short file with a strict grammar and a forgiving reader. Here is the format, the limits, what Google does with each tag, and the difference between an error that costs you one page and one that breaks the whole file under the protocol.
An XML sitemap is a list of the pages on a site that you want search engines to know about, written in the format published at sitemaps.org and read by Google, Bing and the other major engines. It does not make a page rank and it does not force a page into the index. What it does is hand a crawler a complete list of addresses, so discovery stops depending on whether your internal links happen to reach every page. It matters most on a large site, a new site, or one with deep archives that few links reach.
The required structure is small. A root element, <urlset>, which declares the sitemaps.org namespace. Inside it, one <url> element per page. Inside each of those, one <loc> holding the page's full address. That is the entire mandatory grammar. Here is an XML sitemap example that meets it, with the one optional tag worth keeping:
<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<url>
<loc>https://yourdomain.com/</loc>
<lastmod>2026-09-01</lastmod>
</url>
<url>
<loc>https://yourdomain.com/guides/estate-planning-basics</loc>
<lastmod>2026-08-14T09:30:00+00:00</lastmod>
</url>
</urlset>
Three details in that block trip up real files. The namespace is a fixed string, http://www.sitemaps.org/schemas/sitemap/0.9, and it is an identifier rather than a link, so it keeps the http:// even on a site that has been HTTPS for a decade; a generator that "upgrades" it to https:// has written a different namespace. Every <loc> is a fully qualified address with the protocol and host, never a path like /about. And because the file is XML, an ampersand in a URL must be written &. The file must be encoded as UTF-8.
The protocol defines three optional tags per entry. <lastmod> is the date the page last changed, in W3C datetime format: a plain date such as 2026-09-01, or a date and time with a time zone such as 2026-09-01T14:30:00+00:00 or 2026-09-01T14:30Z. <changefreq> is a hint about how often the page changes, from always to never. <priority> is a number from 0.0 to 1.0 describing a page's importance relative to the rest of the site, defaulting to 0.5.
Google's documentation is blunt about two of them: it ignores <priority> and <changefreq>. It uses <lastmod> only if the value is consistently and verifiably accurate, meaning it reflects a real change to the page's main content, structured data or links. A sitemap that stamps every URL with today's date on every build teaches the crawler to stop trusting the tag. So the practical advice is to drop the two ignored tags or leave them as they are, and to spend the effort on making <lastmod> true.
One sitemap file may list at most 50,000 URLs and may be at most 50 MB uncompressed. Both limits are in the sitemaps.org protocol and Google applies the same ones. A file may be served gzipped to save bandwidth, but the 50 MB limit is measured on the unpacked XML.
A site with more pages than that splits its list across several files and publishes a sitemap index, a second kind of file whose root is <sitemapindex> and whose entries are <sitemap> elements, each with the <loc> of one child sitemap. An index can list up to 50,000 child sitemaps under the same 50 MB ceiling. Many sites use an index long before they need one, because content management systems split sitemaps by type: posts, pages, categories, authors. That is fine, and it means the file most people find at /sitemap.xml or /sitemap_index.xml is an index rather than a list of pages.
This validator handles both. Given an index, it fetches the first five children and reports each one's status code, URL count and error count, so a broken child shows up without you opening every file. To see every error in a particular child, enter that child's own URL.
Most validators list errors as equals. They are not. Some make the protocol treat the whole file as invalid; others leave the file readable and cost you one entry, or one entry's date. The table is every error this validator reports, in the words the card uses, with what it means for the pages in the file.
| Error | Effect | What it means for your pages |
|---|---|---|
| Empty file, or an HTML page | File invalid | There is no sitemap at that address. An HTML answer is usually a soft 404 or a redirect to the homepage, and it hides the fact that the file is missing. |
| No <urlset> or <sitemapindex> root | File invalid | The body is not a sitemap: a feed, a text file, a gzip archive served under the wrong name, or a file cut off before its root. |
| Namespace does not match | File invalid under the protocol | The protocol requires its exact namespace on the root element. Often the namespace is present but written with https://, which is a different string. How Google treats that file is not published, and this tool does not predict it. |
| No entries, or over 50,000 | File invalid | An empty list tells a crawler nothing. Past 50,000 URLs the file breaks the protocol and has to be split behind an index. |
| Entries without a <loc> | Entry dropped | The entry has no address, so there is nothing to crawl. The rest of the file is still read. |
| <loc> not an absolute URL | Entry dropped | A relative path such as /about cannot be resolved from a sitemap, so that page is never listed at all. |
| <loc> on another host | Entry dropped | A sitemap speaks for its own host. A www file listing bare-domain pages, or a staging host left in the URLs, falls in this row. |
| <lastmod> not W3C format | Date not used | The date cannot be read, so the page loses the one freshness signal Google uses from a sitemap. 09/12/2026 is ambiguous and fails; 2026-09-12 passes. |
| Duplicate <loc> | Entry dropped | The second copy adds nothing. Duplicates usually point to a generator listing the same page from two sources. |
The comparison here is between a strict reading of the protocol and a lenient one. The protocol is strict, and a file-level error means the file does not meet it. Search engines are more forgiving in practice, and Google does not publish how it handles each defect, which is why the card never tells you Google has rejected or accepted anything. SEOTesting and Nuxt SEO draw the same strict-versus-lenient line on their validators. What this one adds is the consequence stated beside each error, so the person reviewing a vendor's sitemap can tell a cosmetic issue from a missing section of the site. If Search Console reports that it could not read your sitemap, the file-level rows are where to look first.
A sitemap says "please crawl these pages". A robots.txt Disallow rule says "do not crawl these pages". When the same URL appears in both, the site is asking for two opposite things, and the robots.txt wins: a crawler that respects the file will not fetch a page it is told to avoid, whatever the sitemap says. The listing is wasted, and on a site that relies on that page being found, so is the page.
It happens more often than it should. A folder is blocked during a redesign and never unblocked. A plugin generates the sitemap from every published post while a separate rule blocks a category. A migration leaves an old path in the disallow list that now matches live pages. Nothing warns you, because each file is valid on its own. This validator reads the site's robots.txt from the sitemap's own host and tests the first 200 listed URLs against it as Googlebot, using the same longest-match rule Google documents, and names up to five of the blocked addresses.
The other half of the same contradiction is a listed page that carries a noindex tag: the sitemap asks for it to be crawled and the page asks not to be indexed. This tool does not open the listed pages, so it cannot see those tags. The Noindex Checker reads them one URL at a time.
The protocol ties a sitemap to its location: a file at https://yourdomain.com/sitemap.xml can list any page on that host, while a file at https://yourdomain.com/blog/sitemap.xml may only list pages under /blog/. Google relaxes the rule for sitemaps submitted in Search Console, where one file can cover several verified sites. Keeping the sitemap at the root avoids the question. This validator checks the host of each URL but not the folder rule, so a sitemap stored in a subfolder that lists pages above it will pass here.
Crawlers find a sitemap in three ways: a Sitemap: line in robots.txt, a submission in Search Console or Bing Webmaster Tools, or the conventional address. Google announced in June 2023 that its old sitemap ping endpoint was going away, so a tool that offers to ping Google on your behalf is offering something that no longer exists. Whether your site declares a sitemap at all, and where, is a different question from whether the file is well formed; the Sitemap Checker answers that one from your domain.
xmlns="https://www.sitemaps.org/..." looks modern and does not match the protocol. The identifier is the http:// string, exactly./pricing instead of the full address. Every <loc> needs the protocol and host.www.yourdomain.com listing yourdomain.com pages, or the other way round, after the preferred host changed. Pick the host your canonical tags use and list only that one.12/09/2026 means different days in different countries, which is why the protocol requires the W3C form.& breaks strict XML parsers. Write &, or keep query strings out of the sitemap entirely.A root <urlset> with the exact namespace. One entry for each canonical, indexable page, with its full address on one host. A <lastmod> that changes only when the page does. Under 50,000 entries per file, with an index once you pass that. Declared in robots.txt with a Sitemap: line, and never listing a path the same robots.txt blocks. Then re-check after every migration, plugin change or redesign, because sitemaps are generated files and generators change their minds.
Six things that are true of this tool, each one backed by a line in the code that runs it.
The request carries one URL and nothing else. There is no account, no session and no database behind the tool, so there is nothing for us to keep about you.
Each error carries its effect: the file is invalid under the protocol, or the file is still read and that entry is dropped. You can tell a cosmetic defect from a missing section of the site at a glance.
Up to 200 listed URLs are tested against the site's robots.txt as Googlebot, and up to five blocked addresses are named, so a page you list and forbid in the same breath shows up.
Enter an index and the first five child sitemaps are fetched and summarised with their status, URL count and error count, instead of stopping at the list of file names.
Fixed rules, no model, so the same file gives the same answer every time. What it does not check, from XML well-formedness to noindex tags, is listed beside the result rather than in a footnote.
A run is a handful of small fetches, so it costs nothing and is never metered. The one limit is a courtesy to the sites being fetched: 20 checks a minute per visitor.
What breaks a sitemap, what Google ignores, and the limits of one file.
Under the sitemaps.org protocol, a file is invalid when it has no <urlset> or <sitemapindex> root, when the root lacks the exact namespace http://www.sitemaps.org/schemas/sitemap/0.9, when it has no entries, or when it goes past 50,000 URLs or 50 MB uncompressed. An empty body or an HTML page at the sitemap's address also means there is no valid file. Entry-level defects, such as a relative or cross-host <loc>, a missing <loc>, a badly formatted <lastmod> or a duplicate, leave the file readable and mean only that entry is not used. The validator labels every error with one effect or the other.
Google does not publish how it treats each defect, so no outside tool can promise either way, and this one never says Google has accepted or rejected a file. What can be said: the protocol is strict and search engines read more leniently, so an entry-level error usually costs you that entry rather than the file. File-level errors are the serious ones. If Search Console says your sitemap could not be read, check the file-level rows first: the status code, the root element and the namespace. Search Console's Sitemaps report is where Google's own reading of the file shows.
A sitemap index is a file whose root is <sitemapindex> and whose entries point at other sitemap files, used when a site has more than 50,000 URLs or when a content management system splits sitemaps by type. When you enter an index, the validator checks the index itself and then fetches its first five children, reporting each one's status code, URL count and error count. It does not list every error inside a child, and it does not test child URLs against robots.txt. For the full report on any child, enter that child's own URL.
Not to Google, which documents that it ignores both <priority> and <changefreq>. They are optional in the protocol, so leaving them out does not make a file invalid, and this validator does not check their values. Other search engines may read them as hints. The optional tag Google does use is <lastmod>, and only when it is consistently accurate, so that is the one worth maintaining.
W3C datetime format. The validator accepts a plain date, 2026-09-12, or a date with a time and a time zone: 2026-09-12T14:30Z, 2026-09-12T14:30:00+01:00, with optional fractional seconds. A time without a time zone, a date written as 09/12/2026, or text such as "yesterday" fails the check. The format is only half of it: Google uses the value only when it reflects a real change to the page, so a date that updates on every build is valid and ignored.
50,000 URLs and 50 MB uncompressed per file, from the sitemaps.org protocol, which Google applies as well. A sitemap index has the same 50 MB ceiling and can list up to 50,000 child sitemaps. Gzip compression is allowed, but the limit is measured on the unpacked file. This validator counts entries against the 50,000 limit and reports the file's size. It reads at most 8 MB of any file, so on a larger one it says the count covers only the first 8 MB and does not judge the 50 MB limit.
A valid file is the list, not the content on it. Verand writes articles from your own expertise and credentials, so what the crawlers find is yours. It then tracks where you rank on Google and where ChatGPT, Gemini, Google AI Overviews, Google AI Mode, Perplexity and Claude name you, and gates every draft so a claim your regulator would not allow never publishes.
Content built to rank in
Google and get cited by
ChatGPT
Perplexity
Gemini
Claude, with every claim checked before it goes live.