Enter a domain and get a sitemap.xml built from a crawl of your site: up to 150 pages, starting at the homepage and following its links. A page makes the file only if it answers 200, is indexable, names itself as canonical and is open to Googlebot. Every page left out is counted, with the reason. Already have the list? Paste it and build the file in your browser.
Up to 150 pages, six at a time, each allowed 12 seconds. The crawl stops fetching at 40 seconds, and the file arrives in one piece when it ends, usually well inside a minute.
Up to 150 pages, one host. Breadth-first from the homepage, following every same-site link in each page's served HTML, six pages at a time as VerandBot/1.0, 12 second timeout, no JavaScript run. robots.txt is read first: a path it disallows for our crawler or for Googlebot is never requested. Your existing sitemap is not read, so the file shows what your links actually reach.
A page is listed only if it answered 200 at the exact address linked, served HTML, has no noindex in its robots meta tag or X-Robots-Tag header, and has no canonical pointing to another URL. Every other page is counted under the reason it was left out, with up to five examples each. A page that answered 429 or 503, or didn't answer, is marked not checked rather than missing, and kept out of the file. Each entry is a <loc> and nothing else, deduplicated, in the order the pages were reached.
Go past 150 pages or 40 seconds: the card says which limit ended the crawl and how many links were left unvisited. Follow links JavaScript adds after load, or find a page nothing links to. Write lastmod, priority or changefreq (on purpose), image, video or hreflang entries, or a sitemap index. Check the URLs you paste. Upload the file, or decide what Google indexes.
Read the rules, walk the links, leave out every page that should not be there, write the rest down. The card is the tool in motion on an example site, looped, and each step lights up while the card is doing it.
Before any page is requested, the tool reads your robots.txt. A path it disallows for Googlebot never makes the file, and a path it disallows for our own crawler is never fetched. A Crawl-delay for us slows the crawl to one page at a time.
Six pages at a time, breadth-first: the homepage, then every same-site link it serves, then theirs. Your existing sitemap is ignored on purpose, since reading it would hand you back the file you already have. Links to images, PDFs and other files are skipped.
Each page is read with the same parser the product's site crawl uses. A redirect, an error status, a noindex or a canonical pointing elsewhere takes a page out of the file, and the card counts it under that reason with examples.
The pages that passed are written as <loc> entries, with no dates or priorities. Copy it or download sitemap.xml, put it at your site's root and point robots.txt and Search Console at it.
Most generators list every URL they reach and decorate each one with a date and a priority. This one lists fewer pages and says why. Here is what goes in a sitemap, what doesn't, why this file carries no dates, and how to put it live without a plugin.
An XML sitemap is a file, usually at https://yourdomain.com/sitemap.xml, that lists the pages on your site you want search engines to know about. It exists for discovery: a crawler that reads it learns about pages it might not reach by following links, such as a new article nothing links to yet or a page three menus deep. It does not rank anything, and it does not get a page indexed. Google's own documentation calls submitting one "merely a hint" that "doesn't guarantee that Google will download the sitemap or use the sitemap for crawling URLs on the site" (Build and submit a sitemap, updated 8 July 2026).
That makes the file a statement more than a lever. Every URL in it says "this is a page I want in search". When the statement is wrong, and the file lists pages that redirect, return errors or ask not to be indexed, it teaches the crawler that your list can't be taken at face value. The value of a sitemap is in its accuracy, which is why this generator spends its effort on what to leave out.
The tool has two modes, because there are two situations.
<loc>. Nothing is fetched, so nothing is checked: the file is only as right as the list you pasted.One design choice in crawl mode matters more than it looks. The crawl starts at your homepage and ignores any sitemap you already have. A crawler that seeds itself from your existing sitemap and writes a new one hands you back your own file, errors included. Starting from the homepage shows you what your links actually reach, which is also roughly what a search engine finds without a sitemap. A page missing from the result is a page nothing links to, and that's worth knowing on its own.
Google's guidance is to list the canonical URLs you want shown in search and, where the same content lives at several addresses, to pick one. In practice that rules out six kinds of page, and they are the six reasons the card counts:
| Left out when | Why it doesn't belong |
|---|---|
| The address redirects | The sitemap should name where the page lives, not a door to it. The tool drops the linked address and checks the destination as a page of its own, so the real page can still make the file. |
| It doesn't answer 200 | A 404, a 410 or a server error. Listing it asks the crawler to visit a page that isn't there. |
| It carries noindex | In the robots meta tag or the X-Robots-Tag header. A page can't sensibly ask to be found and to be left out of the index at once. |
| Its canonical points elsewhere | The page says another URL is the real one. List that one instead; this tool lists it when the crawl reaches it. |
| robots.txt blocks it | Disallowed for Googlebot, so Google won't fetch it. Listing a blocked page is a contradiction the Sitemap Validator flags in existing files. |
| It isn't HTML | The response had no page to read. Images, PDFs and similar files are skipped before they're fetched. |
One answer gets its own label instead of a place in that table. A 429 means the site is asking the crawler to slow down, a 503 that it's temporarily unavailable, and no answer at all within 12 seconds usually means a slow or busy server. None of them says the page is missing, so the card marks those pages not checked, keeps them out of the file because they couldn't be verified, and tells you to run it again later. The tool doesn't retry them in the same run: re-asking a site that just said slow down, at the pace that provoked it, is the behaviour the 429 objects to.
Two edge cases follow from reading URLs exactly. /services and /services/ are different addresses, and on most sites one redirects to the other, so only one of them is listed. And a URL with a query string, such as a sorted or filtered view, is listed like any other page if it answers 200 and names itself as canonical. If your filtered pages do that, the fix is their canonical tag, not the sitemap.
The sitemap protocol allows three optional tags per URL. Google's documentation settles two of them in one sentence: "Google ignores <priority> and <changefreq> values." The third, <lastmod>, it uses only "if it's consistently and verifiably (for example by comparing to the last modification of the page) accurate."
A crawler can't know when your content last changed. The dates generators write come from a server header, the time of the crawl, or nothing at all, and a date that moves every time the file is rebuilt is exactly the inaccuracy Google describes. Priorities set by URL depth are guesses about your business written in your name. So this generator writes neither, the same rule Verand's product follows: a wrong date is worse than none. If your CMS knows the real modification time of each page, let it write lastmod; that is information only your site has.
This is the start of the file the tool wrote for verand.ai on 29 September 2026. The crawl reached every linked page, 130 of them, in under ten seconds, and listed 125:
<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<url><loc>https://verand.ai/</loc></url>
<url><loc>https://verand.ai/pricing</loc></url>
<url><loc>https://verand.ai/features/article-generator</loc></url>
[122 more entries]
</urlset>
The five pages it left out are the useful part. All five answered 404, and all five are links on our own site to pages that haven't been built yet, such as /insurance-agencies/advertising-rules. A generator that lists everything it reaches would have put them in the sitemap. This one kept them out and named them, which turns the sitemap run into a broken-link report for the pages it read.
The same day, willowdaleequity.com, a site with roughly 245 articles, filled all 150 pages with nothing to leave out, and the card said the crawl stopped at the cap with 99 same-site links still unvisited. That result is a partial list, and the card says so rather than presenting it as the whole site.
The file is never padded and never silently partial. The card reports which of these ended the crawl, and how many links it found but didn't visit:
Crawl-delay; we honour it because it's addressed to us.)Usually not, and the honest answer is to check before you replace anything. WordPress has served a basic sitemap at /wp-sitemap.xml since version 5.5, and SEO plugins such as Yoast and Rank Math swap in their own. Shopify, Wix and Squarespace generate /sitemap.xml for every site automatically. Static site frameworks have sitemap plugins or conventions, such as @astrojs/sitemap for Astro or a sitemap route in Next.js. A file your platform maintains updates itself as you publish, which a file you upload by hand never will.
So the better first step on those platforms is the Sitemap Checker, which finds the file you already serve, then the Sitemap Validator, which reads it for errors. Where this generator earns its place: a hand-built or legacy site with no sitemap, a platform whose generated file lists pages it shouldn't, or a quick independent crawl to compare with what your CMS claims.
sitemap.xml, UTF-8, exactly as downloaded. Don't add dates by hand unless they're true.https://yourdomain.com/sitemap.xml. On a static site that means the folder that becomes the root of the build, usually public/ or static/. A sitemap covers the URLs at or below its own location, so the root covers everything.Sitemap: https://yourdomain.com/sitemap.xml. Crawlers that read robots.txt find it there without being told.www.yourdomain.com should list www URLs. Paste mode keeps the lines on the host of your first URL and sets the rest aside for that reason.The file this tool writes is honest about being partial, but it is partial. For a larger site the right source is almost always the platform's own sitemap, which knows every page. Other free crawlers go further than 150 pages, and the Sitemap Validator will tell you whether what they produce lists pages it shouldn't. Past 50,000 URLs, or 50 MB uncompressed, the protocol requires splitting the list across several files joined by a sitemap index, which this tool does not write.
Six things that are true of this tool, each one backed by a line in the code that runs it.
The request carries a domain and nothing else, and paste mode sends nothing at all. There is no account, no session and no database behind the tool, so there is nothing for us to keep about you.
A page is listed only if it answered 200, isn't redirected, carries no noindex in its meta tag or header, has no canonical pointing elsewhere and isn't blocked for Googlebot. The rest are left out, not listed and hoped for.
Counts for six reasons, up to five example URLs each with the detail (the status code, the redirect target, where the noindex was found), pages the site wouldn't let us check marked as such rather than missing, plus which limit ended the crawl and how many links went unvisited.
The file is <loc> only. It's written by the same function Verand's product uses when a customer's site needs a sitemap, which leaves lastmod out because a wrong date is worse than none.
Fixed rules, no model, no paid rendering. It obeys your robots.txt and your Crawl-delay, and every request counts against a per-site budget shared by all our free tools, so it can't be used to hammer a site.
A run costs nothing, so it's never metered. The limits are courtesies: four crawls a minute per visitor, and the per-site budget above. Paste mode has no limit at all, because it runs on your machine.
What goes in the file, what stays out, and what to do with it once it's written.
No. Google's sitemap documentation says it ignores the priority and changefreq values. They are optional in the sitemap protocol, so leaving them out doesn't make a file invalid, and other search engines may read them only as hints. This generator doesn't write them. The one optional tag Google does use is lastmod, and only when it is consistently and verifiably accurate.
Because a crawler can't know when your content last changed. Generators that fill lastmod take it from a server header or the time of the crawl, and a date that moves on every rebuild is the inaccuracy that makes Google stop trusting the tag. A wrong date is worse than none. If your CMS records the real modification date of each page, its own sitemap can carry lastmod truthfully.
Pages that redirect, pages that return an error or don't answer, pages with noindex in the robots meta tag or the X-Robots-Tag header, pages whose canonical points to another URL, and pages robots.txt blocks for Googlebot. Google's guidance is to list the canonical URLs you want in search. This generator leaves out all five kinds, plus anything that isn't an HTML page, and counts each under its reason with examples.
Put it at the root of your site so it answers at https://yourdomain.com/sitemap.xml. Then add one line to robots.txt, Sitemap: https://yourdomain.com/sitemap.xml, and submit the same URL in Google Search Console under Sitemaps. There is no separate sitemap generator for Google: Google reads the standard sitemaps.org format this tool writes. The Sitemap Checker confirms the file can be found once it's live.
Usually not. WordPress, Shopify, Wix and Squarespace all serve a sitemap automatically, and a file your platform maintains stays current as you publish. Check yours with the Sitemap Checker and the Sitemap Validator instead of replacing it. Use this generator for a site with no sitemap, or to compare an independent crawl with what your platform's file lists.
The crawl stops at 150 pages, and the card says so and counts the links it didn't visit, so the file is a partial list. For a larger site, use your platform's sitemap, or paste a full URL list from an export into paste mode, which formats up to 50,000 URLs in your browser. Beyond 50,000 URLs or 50 MB, the protocol calls for several files joined by a sitemap index, which this tool doesn't write.
Verand writes from your own experience, gets you found on Google and in AI answers, and tracks where you rank and which answers name you. Every draft passes a compliance gate before anyone can publish it: your industry's rules if you have them, a truth-in-advertising check if you don't.
Researched from the regulators' own text and tested by Verand. Not reviewed by a licensed attorney. Your counsel confirms applicability. Not legal advice. Example shown is illustrative.
The 2019 fund returned 14.2% net to investors, and this offering carries guaranteed returns of 12 to 15% over a five year hold.
Performance advertising can’t promise a return. The claim needs a basis and the required disclosures, and it still can’t be stated as a guarantee.
Willowdale Equity and four competitors · 14 days
Content built to rank in
Google and get cited by
ChatGPT
Perplexity
Gemini
Claude, with every claim checked before it goes live.