Free SEO Tool · No Signup Required

Free Sitemap Validator

Enter a sitemap's address, or just your domain to check its /sitemap.xml. The validator fetches the file, checks it against the sitemaps.org protocol, names every error and says whether it makes the whole file invalid or only drops one entry, then counts listed URLs your own robots.txt tells Google not to crawl.

https://

Read-only. We fetch the file, its robots.txt and, for an index, up to five child sitemaps. A domain on its own is checked at /sitemap.xml; if yours lives elsewhere, the Sitemap Checker finds it.

  • No signup, no email wall
  • Same result every run
  • Names each protocol error
  • Reads index children
  • Flags robots.txt conflicts
  • Free, no daily cap
What we checked

Up to 10 checks on one file. One GET for the URL you enter, as VerandBot/1.0, redirects followed, 10 second timeout. The root element and its sitemaps.org 0.9 namespace; entries present and under 50,000 URLs; every entry has a <loc> that is absolute and on the sitemap's own host; <lastmod> in W3C format; no duplicate <loc>. Each error is marked as one that invalidates the file or one that drops an entry.

Indexes and robots.txt

For a sitemap index, the first five child sitemaps are fetched and each is summarised with its URL count and error count. For a URL set, the first 200 listed URLs are tested against the site's own robots.txt as Googlebot: the longest matching rule wins, and Allow beats Disallow on a tie.

What it cannot see

The file is read by pattern matching, not a full XML parser, so a malformed tag or an unescaped & can pass. It reads the first 8 MB, so the 50 MB limit is not measured on larger files. Gzipped .xml.gz files are not unpacked. <changefreq>, <priority>, image and video tags, the UTF-8 encoding and noindex tags on the listed pages are not checked. A pass means the file meets these checks, not that Google has accepted it.

About this tool

How the XML sitemap validator reads your file.

One fetch, one pass over the entries, one question per rule. The card is the tool in motion on an example file with two mistakes in it, looped, and each step lights up while the card is doing it.

01

One request for the file you name

The tool asks for the URL you entered, or for /sitemap.xml when you enter only a domain, identifying itself as VerandBot, following redirects, giving up after ten seconds. Anything but a 200 stops here with the status code, because there is nothing to validate.

02

The root decides what the file is

A <urlset> root makes it a list of pages; a <sitemapindex> root makes it a list of sitemaps. An HTML page, an empty body or neither root ends the run with one file-level error.

03

Every entry, every rule

Each entry is read for its <loc> and <lastmod>, and the errors are counted per rule. Then up to 200 of the listed URLs are matched against the site's robots.txt, the way Googlebot would read it.

04

A consequence, not just a code

Every error on the card says whether the protocol treats the whole file as invalid or whether the file is still read and only that entry is dropped. Where a fix is one line, the card prints it.

Sitemaps, explained

What a valid sitemap contains, and which errors actually cost you.

A sitemap is a short file with a strict grammar and a forgiving reader. Here is the format, the limits, what Google does with each tag, and the difference between an error that costs you one page and one that breaks the whole file under the protocol.

Sitemap XML format: the parts that are required

An XML sitemap is a list of the pages on a site that you want search engines to know about, written in the format published at sitemaps.org and read by Google, Bing and the other major engines. It does not make a page rank and it does not force a page into the index. What it does is hand a crawler a complete list of addresses, so discovery stops depending on whether your internal links happen to reach every page. It matters most on a large site, a new site, or one with deep archives that few links reach.

The required structure is small. A root element, <urlset>, which declares the sitemaps.org namespace. Inside it, one <url> element per page. Inside each of those, one <loc> holding the page's full address. That is the entire mandatory grammar. Here is an XML sitemap example that meets it, with the one optional tag worth keeping:

<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
  <url>
    <loc>https://yourdomain.com/</loc>
    <lastmod>2026-09-01</lastmod>
  </url>
  <url>
    <loc>https://yourdomain.com/guides/estate-planning-basics</loc>
    <lastmod>2026-08-14T09:30:00+00:00</lastmod>
  </url>
</urlset>

Three details in that block trip up real files. The namespace is a fixed string, http://www.sitemaps.org/schemas/sitemap/0.9, and it is an identifier rather than a link, so it keeps the http:// even on a site that has been HTTPS for a decade; a generator that "upgrades" it to https:// has written a different namespace. Every <loc> is a fully qualified address with the protocol and host, never a path like /about. And because the file is XML, an ampersand in a URL must be written &amp;. The file must be encoded as UTF-8.

The optional tags, and what Google does with them

The protocol defines three optional tags per entry. <lastmod> is the date the page last changed, in W3C datetime format: a plain date such as 2026-09-01, or a date and time with a time zone such as 2026-09-01T14:30:00+00:00 or 2026-09-01T14:30Z. <changefreq> is a hint about how often the page changes, from always to never. <priority> is a number from 0.0 to 1.0 describing a page's importance relative to the rest of the site, defaulting to 0.5.

Google's documentation is blunt about two of them: it ignores <priority> and <changefreq>. It uses <lastmod> only if the value is consistently and verifiably accurate, meaning it reflects a real change to the page's main content, structured data or links. A sitemap that stamps every URL with today's date on every build teaches the crawler to stop trusting the tag. So the practical advice is to drop the two ignored tags or leave them as they are, and to spend the effort on making <lastmod> true.

Size limits, and the sitemap index

One sitemap file may list at most 50,000 URLs and may be at most 50 MB uncompressed. Both limits are in the sitemaps.org protocol and Google applies the same ones. A file may be served gzipped to save bandwidth, but the 50 MB limit is measured on the unpacked XML.

A site with more pages than that splits its list across several files and publishes a sitemap index, a second kind of file whose root is <sitemapindex> and whose entries are <sitemap> elements, each with the <loc> of one child sitemap. An index can list up to 50,000 child sitemaps under the same 50 MB ceiling. Many sites use an index long before they need one, because content management systems split sitemaps by type: posts, pages, categories, authors. That is fine, and it means the file most people find at /sitemap.xml or /sitemap_index.xml is an index rather than a list of pages.

This validator handles both. Given an index, it fetches the first five children and reports each one's status code, URL count and error count, so a broken child shows up without you opening every file. To see every error in a particular child, enter that child's own URL.

Sitemap errors, and what each one costs

Most validators list errors as equals. They are not. Some make the protocol treat the whole file as invalid; others leave the file readable and cost you one entry, or one entry's date. The table is every error this validator reports, in the words the card uses, with what it means for the pages in the file.

ErrorEffectWhat it means for your pages
Empty file, or an HTML pageFile invalidThere is no sitemap at that address. An HTML answer is usually a soft 404 or a redirect to the homepage, and it hides the fact that the file is missing.
No <urlset> or <sitemapindex> rootFile invalidThe body is not a sitemap: a feed, a text file, a gzip archive served under the wrong name, or a file cut off before its root.
Namespace does not matchFile invalid under the protocolThe protocol requires its exact namespace on the root element. Often the namespace is present but written with https://, which is a different string. How Google treats that file is not published, and this tool does not predict it.
No entries, or over 50,000File invalidAn empty list tells a crawler nothing. Past 50,000 URLs the file breaks the protocol and has to be split behind an index.
Entries without a <loc>Entry droppedThe entry has no address, so there is nothing to crawl. The rest of the file is still read.
<loc> not an absolute URLEntry droppedA relative path such as /about cannot be resolved from a sitemap, so that page is never listed at all.
<loc> on another hostEntry droppedA sitemap speaks for its own host. A www file listing bare-domain pages, or a staging host left in the URLs, falls in this row.
<lastmod> not W3C formatDate not usedThe date cannot be read, so the page loses the one freshness signal Google uses from a sitemap. 09/12/2026 is ambiguous and fails; 2026-09-12 passes.
Duplicate <loc>Entry droppedThe second copy adds nothing. Duplicates usually point to a generator listing the same page from two sources.

The comparison here is between a strict reading of the protocol and a lenient one. The protocol is strict, and a file-level error means the file does not meet it. Search engines are more forgiving in practice, and Google does not publish how it handles each defect, which is why the card never tells you Google has rejected or accepted anything. SEOTesting and Nuxt SEO draw the same strict-versus-lenient line on their validators. What this one adds is the consequence stated beside each error, so the person reviewing a vendor's sitemap can tell a cosmetic issue from a missing section of the site. If Search Console reports that it could not read your sitemap, the file-level rows are where to look first.

The contradiction: listing a page and blocking it

A sitemap says "please crawl these pages". A robots.txt Disallow rule says "do not crawl these pages". When the same URL appears in both, the site is asking for two opposite things, and the robots.txt wins: a crawler that respects the file will not fetch a page it is told to avoid, whatever the sitemap says. The listing is wasted, and on a site that relies on that page being found, so is the page.

It happens more often than it should. A folder is blocked during a redesign and never unblocked. A plugin generates the sitemap from every published post while a separate rule blocks a category. A migration leaves an old path in the disallow list that now matches live pages. Nothing warns you, because each file is valid on its own. This validator reads the site's robots.txt from the sitemap's own host and tests the first 200 listed URLs against it as Googlebot, using the same longest-match rule Google documents, and names up to five of the blocked addresses.

The other half of the same contradiction is a listed page that carries a noindex tag: the sitemap asks for it to be crawled and the page asks not to be indexed. This tool does not open the listed pages, so it cannot see those tags. The Noindex Checker reads them one URL at a time.

Where the file lives, and how crawlers find it

The protocol ties a sitemap to its location: a file at https://yourdomain.com/sitemap.xml can list any page on that host, while a file at https://yourdomain.com/blog/sitemap.xml may only list pages under /blog/. Google relaxes the rule for sitemaps submitted in Search Console, where one file can cover several verified sites. Keeping the sitemap at the root avoids the question. This validator checks the host of each URL but not the folder rule, so a sitemap stored in a subfolder that lists pages above it will pass here.

Crawlers find a sitemap in three ways: a Sitemap: line in robots.txt, a submission in Search Console or Bing Webmaster Tools, or the conventional address. Google announced in June 2023 that its old sitemap ping endpoint was going away, so a tool that offers to ping Google on your behalf is offering something that no longer exists. Whether your site declares a sitemap at all, and where, is a different question from whether the file is well formed; the Sitemap Checker answers that one from your domain.

Common mistakes

  • The https namespace. xmlns="https://www.sitemaps.org/..." looks modern and does not match the protocol. The identifier is the http:// string, exactly.
  • Relative URLs. Hand-written and static-site sitemaps often list /pricing instead of the full address. Every <loc> needs the protocol and host.
  • The wrong host. A sitemap on www.yourdomain.com listing yourdomain.com pages, or the other way round, after the preferred host changed. Pick the host your canonical tags use and list only that one.
  • A soft 404 at /sitemap.xml. The site returns its homepage or a themed "not found" page with status 200. Every tool that does not look at the body reports a sitemap that is not there.
  • Dates in a local format. 12/09/2026 means different days in different countries, which is why the protocol requires the W3C form.
  • A lastmod that is always today. Valid, and useless: when every page claims to have changed on every build, the date stops meaning anything.
  • Listing redirects, noindexed pages or blocked paths. A sitemap should list the canonical, indexable version of each page and nothing else. Anything else sends a crawler somewhere you do not want it to stop.
  • Unescaped ampersands. A query string with a bare & breaks strict XML parsers. Write &amp;, or keep query strings out of the sitemap entirely.

What a good file looks like

A root <urlset> with the exact namespace. One entry for each canonical, indexable page, with its full address on one host. A <lastmod> that changes only when the page does. Under 50,000 entries per file, with an index once you pass that. Declared in robots.txt with a Sitemap: line, and never listing a path the same robots.txt blocks. Then re-check after every migration, plugin change or redesign, because sitemaps are generated files and generators change their minds.

Why this one

Why choose Verand's Sitemap Validator?

Six things that are true of this tool, each one backed by a line in the code that runs it.

No signup, no email wall

The request carries one URL and nothing else. There is no account, no session and no database behind the tool, so there is nothing for us to keep about you.

A consequence beside every error

Each error carries its effect: the file is invalid under the protocol, or the file is still read and that entry is dropped. You can tell a cosmetic defect from a missing section of the site at a glance.

Checks your own robots.txt

Up to 200 listed URLs are tested against the site's robots.txt as Googlebot, and up to five blocked addresses are named, so a page you list and forbid in the same breath shows up.

Follows an index one level down

Enter an index and the first five child sitemaps are fetched and summarised with their status, URL count and error count, instead of stopping at the list of file names.

Deterministic, and honest about it

Fixed rules, no model, so the same file gives the same answer every time. What it does not check, from XML well-formedness to noindex tags, is listed beside the result rather than in a footnote.

$0, no daily cap

A run is a handful of small fetches, so it costs nothing and is never metered. The one limit is a courtesy to the sites being fetched: 20 checks a minute per visitor.

Questions

Frequently Asked Questions About the Sitemap Validator

What breaks a sitemap, what Google ignores, and the limits of one file.

What makes a sitemap invalid?

Under the sitemaps.org protocol, a file is invalid when it has no <urlset> or <sitemapindex> root, when the root lacks the exact namespace http://www.sitemaps.org/schemas/sitemap/0.9, when it has no entries, or when it goes past 50,000 URLs or 50 MB uncompressed. An empty body or an HTML page at the sitemap's address also means there is no valid file. Entry-level defects, such as a relative or cross-host <loc>, a missing <loc>, a badly formatted <lastmod> or a duplicate, leave the file readable and mean only that entry is not used. The validator labels every error with one effect or the other.

Does Google reject a sitemap with minor errors?

Google does not publish how it treats each defect, so no outside tool can promise either way, and this one never says Google has accepted or rejected a file. What can be said: the protocol is strict and search engines read more leniently, so an entry-level error usually costs you that entry rather than the file. File-level errors are the serious ones. If Search Console says your sitemap could not be read, check the file-level rows first: the status code, the root element and the namespace. Search Console's Sitemaps report is where Google's own reading of the file shows.

What is a sitemap index, and are child sitemaps validated too?

A sitemap index is a file whose root is <sitemapindex> and whose entries point at other sitemap files, used when a site has more than 50,000 URLs or when a content management system splits sitemaps by type. When you enter an index, the validator checks the index itself and then fetches its first five children, reporting each one's status code, URL count and error count. It does not list every error inside a child, and it does not test child URLs against robots.txt. For the full report on any child, enter that child's own URL.

Do priority and changefreq matter?

Not to Google, which documents that it ignores both <priority> and <changefreq>. They are optional in the protocol, so leaving them out does not make a file invalid, and this validator does not check their values. Other search engines may read them as hints. The optional tag Google does use is <lastmod>, and only when it is consistently accurate, so that is the one worth maintaining.

What does lastmod need to look like?

W3C datetime format. The validator accepts a plain date, 2026-09-12, or a date with a time and a time zone: 2026-09-12T14:30Z, 2026-09-12T14:30:00+01:00, with optional fractional seconds. A time without a time zone, a date written as 09/12/2026, or text such as "yesterday" fails the check. The format is only half of it: Google uses the value only when it reflects a real change to the page, so a date that updates on every build is valid and ignored.

What are the size limits for one sitemap file?

50,000 URLs and 50 MB uncompressed per file, from the sitemaps.org protocol, which Google applies as well. A sitemap index has the same 50 MB ceiling and can list up to 50,000 child sitemaps. Gzip compression is allowed, but the limit is measured on the unpacked file. This validator counts entries against the 50,000 limit and reports the file's size. It reads at most 8 MB of any file, so on a larger one it says the count covers only the first 8 MB and does not judge the 50 MB limit.

After the check

A clean sitemap gets pages found. Make them worth finding.

A valid file is the list, not the content on it. Verand writes articles from your own expertise and credentials, so what the crawlers find is yours. It then tracks where you rank on Google and where ChatGPT, Gemini, Google AI Overviews, Google AI Mode, Perplexity and Claude name you, and gates every draft so a claim your regulator would not allow never publishes.

Verand

Content built to rank in Google and get cited by ChatGPTPerplexityGeminiClaude, with every claim checked before it goes live.

support@verand.ai

© 2026 Verand. All rights reserved. TermsPrivacyAI policyAccessibilitySecurity
Not legal advice. Compliance packs are researched from the regulators' own text and tested by Verand, not reviewed by a licensed attorney.