Free SEO Tool · No Signup Required

Free Robots.txt Generator

Build a robots.txt from a preset or from the file your site already serves. Set rules for Googlebot and for GPTBot, ClaudeBot and 12 other AI crawlers by name, and every group you add keeps the rules you wrote for everyone else. The file updates as you click, in your browser.

https://

Optional. We read your live robots.txt and look for your sitemap, then load them below so you edit your file instead of overwriting it. Nothing is written to your site.

  • No signup, no email wall
  • Starts from your file
  • Keeps your * rules
  • 14 AI crawlers by name
  • Live preview in-browser
  • Free, no daily cap
New robots.txtStarting point: the Allow all preset
Ready to review
Start fromA preset replaces the rules for everyone else. Your crawler choices stay.
Everyone else User-agent: *Every crawler you don't name below follows these lines.
AI crawlersFollow * writes nothing. Allow writes a group that repeats your * rules. Block writes Disallow: /.

For what each crawler does and a policy by purpose, use the Robots.txt Generator for AI.

Search engines and other crawlersName any crawler by its user-agent token.
SitemapFull address, one per line.
robots.txt2 lines

        
Re-readYour file, parsed again the way a crawler groups it: longest matching rule wins, Allow wins a tie.

After uploading it to the root of your site, confirm what went live in the Robots.txt Checker.

What it writes

A User-agent: * group from a preset or your own rules, a group for each crawler you set to Allow or Block, an optional Crawl-delay and your Sitemap lines. Built in your browser; only "Start from my file" calls our server, which reads /robots.txt and probes for a sitemap.

How your * rules carry over

A crawler that has its own group ignores User-agent: * entirely, so every Allow group repeats your * lines and Crawl-delay. A crawler your live file already names is left exactly as written, with a Replace button if you mean to change it.

What it can't do

The file is a request crawlers choose to honour: not security, and not a way to keep a page out of Google (that's noindex). Google ignores Crawl-delay. The AI list is the 14 crawlers Verand's audit watches, not every bot. Review the file before you upload it.

About this tool

How this robots.txt file generator writes your file.

Start from what you have, decide per crawler, and let the precedence rule do its work in the open. The card is the tool in motion on an example file, looped, and each step lights up while the card is doing it.

01

Start from what you have

Type a domain and the tool reads its live /robots.txt through our server, or pick a preset. The rules for everyone else, the Crawl-delay and the Sitemap lines land in the builder; groups the file already has for named crawlers are kept as written.

02

Decide per crawler

Each crawler is Follow *, Allow or Block. Block writes a two-line group with Disallow: /. The file on the right rebuilds on every click, in the page, with no request to anyone.

03

Your * rules travel with it

A named crawler ignores User-agent: *, so an Allow group that said only Allow: / would open /admin/ to it. The generator copies every * line into the group instead, and a comment in the file says why.

04

Re-read, then copy

The finished file is parsed again and every crawler is judged on the path you choose, so you see what a crawler would do before anything goes live. Copy or download it, upload it to your site's root, and check it there.

Writing robots.txt

Robots.txt examples, and what each line tells a crawler.

A robots.txt is short, public and read before anything else on your site. Here is where it goes, the five lines a generator writes, the presets and when each one is right, and the one precedence rule that quietly undoes most hand-edited files.

Where the file lives

One file, at one address: https://yourdomain.com/robots.txt. Crawlers only ever request it at the root of a host, so a copy in /blog/robots.txt is never read, and a rule in it does nothing. Google puts the scope precisely: the rules "apply only to the host, protocol, and port number where the robots.txt file is hosted." So shop.yourdomain.com needs its own file, and so does the http:// version of a site if it still answers on its own rather than redirecting.

Two practical limits. Google reads the first 500 KiB and ignores the rest, which no hand-written file comes near but a generated list of thousands of paths can. And Google "generally caches the contents of robots.txt file for up to 24 hours", so a change is not live for Google the moment you upload it. Plan edits around that, particularly the ones that open something up.

The five lines a generator writes

Robots.txt syntax is a list of field: value lines. Field names are not case-sensitive; paths are, so /Admin/ and /admin/ are different rules. Comments start with #. Everything this generator produces is built from five fields:

  • User-agent opens a group and names the crawler it speaks to. * means every crawler that has no group of its own. Consecutive User-agent lines share one set of rules.
  • Disallow is a path prefix the crawler should not request. Disallow: /cart covers /cart, /cart/ and /cartoons, because it is a prefix. An empty Disallow: blocks nothing.
  • Allow carves an exception out of a Disallow. When two rules match a URL, the longer one wins, and on a tie Allow wins.
  • Sitemap points crawlers at your XML sitemap. It belongs to no group and can sit anywhere, but it must be the full address: Google requires "a fully qualified URL, including the protocol and host".
  • Crawl-delay asks a crawler to wait that many seconds between requests. It is not part of the standard, and support is uneven (more below).

Two wildcard characters work in paths: * matches any run of characters and $ anchors the end of the URL. They are defined in RFC 9309, the IETF standard the protocol finally got in 2022, and Google supports both. A robots.txt wildcard is the difference between blocking every PDF and blocking one folder:

# every URL ending in .pdf
Disallow: /*.pdf$

# any URL carrying a session parameter
Disallow: /*?sessionid=

# the folder, but not the one file inside it the page needs
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php

Robots.txt examples for the common cases

The presets in the builder are the four starting points almost every site needs. Each is a sensible base to edit, not a finished policy.

Allow all. Two lines, User-agent: * and Allow: /. It is what robots.txt allow all looks like when written down, and it behaves exactly like having no file: a crawler with no restriction may fetch everything. Writing it anyway gives you a place to put the Sitemap line and a record that the openness is deliberate.

User-agent: *
Allow: /

Sitemap: https://yourdomain.com/sitemap.xml

WordPress. WordPress serves a virtual robots.txt of its own when no real file exists: Disallow: /wp-admin/ with an exception for admin-ajax.php, which themes and plugins call from the public site. Used as a robots.txt generator for WordPress, the preset writes exactly those two rules. It does not block /wp-includes/ or /wp-content/, because both hold the CSS and JavaScript your pages load, and Google renders pages with them. If an SEO plugin already manages the file for you, start from your live file instead, so the plugin's lines are what you edit.

Block everything. Disallow: / under *, the staging preset. More on it next, because it is the most useful and the most dangerous line in the protocol.

My live file. Appears after "Start from my file" and restores exactly what your site served, so you can always return to it after trying a preset.

Robots.txt disallow all: when it's right, and what it doesn't do

User-agent: * followed by Disallow: / asks every crawler that honours the file to request nothing. It is right for a staging copy, a development host, or a site that is not ready to be seen. It is wrong everywhere else, and the usual way it ends up on a live site is a staging file copied to production along with everything else. It stops Googlebot too, which is why the builder flags it every time it's in the output.

What it does not do matters as much. It is not a lock: the file is public and compliance is voluntary, so a staging site that must stay private needs a password. And it does not remove pages from Google. A blocked URL can still be indexed if other pages link to it; Google shows the bare address, because it was told not to read the page. Keeping a page out of the index is the job of a noindex meta tag or X-Robots-Tag header, which the crawler has to be allowed to fetch before it can see. Blocking a page and asking for it to be de-indexed is the classic way to get neither, and Google stopped reading noindex written inside robots.txt in 2019.

Why a named group has to repeat your * rules

This is the rule most generated files get wrong, and the reason this generator writes groups the way it does. Under RFC 9309 a crawler obeys the group that names it and ignores User-agent: * entirely. It does not merge the two. So this file, which looks like "keep everyone out of the client portal, and make sure GPTBot is welcome", does something else:

User-agent: *
Disallow: /client-portal/

User-agent: GPTBot
Allow: /     # GPTBot never reads the * group, so /client-portal/ is open to it

Every tool that adds a per-crawler group from a dropdown can produce that file. This one copies each line of your * group into every Allow group it writes, Allow and Disallow alike, plus the Crawl-delay, and adds a comment to the file saying so, because the copy only holds while the two stay in step. Change the * rules later and the named groups need the same change. Block groups are the exception: they are Disallow: / alone, since an Allow line copied into a block would reopen exactly the path it names.

A crawler your live file already names is treated differently. A group for a specific crawler is the one explicit statement of intent a robots.txt can make: counsel's decision, a CDN setting, a previous agency's policy. Writing a second group for it would be merged with the first and could reverse it. So the generator keeps that group exactly as written, comments included, and shows the crawler as "In your file". The Replace button beside it is there when you do mean to change it. It's the same rule Verand's own fix for blocked AI crawlers follows on customer sites.

Robots.txt crawl-delay: who reads it

Crawl-delay is the line most generators offer and fewest crawlers use. Google's documentation lists the fields it supports and says plainly that "other fields such as crawl-delay aren't supported." Anthropic's says the opposite for its crawlers: "we support the non-standard Crawl-delay extension to robots.txt." So a delay can slow some crawlers and will never slow Googlebot. If Google is putting real load on your server, its documentation points to answering with a 503 or 429 status for a while, not to robots.txt. One more reason the generator copies the delay into each Allow group: a named crawler that honours Crawl-delay only reads it from its own group.

AI crawlers are one section of the file, not a separate file

The AI crawlers in the builder are the 14 names Verand's site audit watches on every customer site, read from the product's own list when this page is built, so the two can't drift apart. They do different jobs: some collect pages for model training, some build the index an assistant searches, and some fetch a page because a person just asked a question. That split decides what a block achieves. OpenAI says of its user-triggered fetcher that "robots.txt rules may not apply", and Perplexity says its own "generally ignores robots.txt rules". Google-Extended is a control token rather than a crawler, and Google says it "does not impact a site's inclusion in Google Search". The Block training crawlers button sets the five the list classes as training and leaves the rest alone; for the purpose-by-purpose policy, the Robots.txt Generator for AI goes crawler by crawler.

Lines never to write

  • Disallow on the CSS or JavaScript a page needs. Google renders pages with their scripts and styles. Block /assets/ or *.js and Google reads a broken page. The builder warns on rules that look like it.
  • A Disallow on the page you're trying to noindex. The crawler can then never fetch the page, so it never sees the noindex. Allow the crawl, serve the tag, and remove the page from your sitemap.
  • A relative Sitemap line. Sitemap: /sitemap.xml is not a full address, and Google requires one. The builder flags any line that doesn't start with http.
  • A path without its leading slash. Disallow: admin is not a path. Rules start with / or a wildcard.
  • A rule above the first User-agent line. It belongs to no group and RFC 9309 says to ignore it. When you start from your file, lines like that are listed as not carried over rather than silently lost.
  • A misspelled crawler name. Names match without regard to case, but GPT-Bot is not GPTBot. A group for a crawler that doesn't exist blocks nothing and looks like it does.

Start from the file you have

Most sites already serve a robots.txt, and most of those files contain a line somebody added for a reason: a search results path, a staging folder, an old campaign directory. A generator that starts from a blank form invites you to overwrite all of it. "Start from my file" reads what your site serves today and loads it: the * rules and Crawl-delay into the builder, the Sitemap lines into the sitemap box, and every group for a named crawler kept word for word. If the file declares no sitemap and one answers at a standard address, that address is filled in and the note says it was found, not declared. Comments outside those named groups, and any line the parser can't read, are not carried over, and the note under the input lists what was left out. If the file can't be read at all, because of a server error or a response that isn't a text file, nothing is imported and the note says so, rather than presenting part of a file as the whole.

Before you upload

Read the output once, top to bottom, with your own site in mind: the generator knows the rules of the format, not the reasons behind your paths. Save it as plain text named robots.txt, put it at the root of the host, and load https://yourdomain.com/robots.txt in a browser to see that it's served as text and not as an HTML page (a theme's 404 page returned as a 200 is read as no file at all). Then run the Robots.txt Checker against the live site. The file is a set of requests to crawlers that choose to honour them; it controls crawling, not rankings, and nothing about it guarantees what any engine does with your pages.

Why this one

Why choose Verand's Robots.txt Generator?

Six things that are true of this tool, each one backed by the code that runs it.

Starts from your live file

One click reads the robots.txt your site serves today and loads it into the builder, so the lines your developer added are what you edit, not what you overwrite.

Keeps your * rules in every group

Each Allow group repeats your * lines and Crawl-delay, and a crawler your file already names is left as written. The same rules as the fix Verand ships for blocked AI crawlers.

Re-reads its own output

The finished file is parsed again with Verand's robots.txt parser, and every crawler is judged on the path you type, with the rule that decided it. You see the outcome before you upload.

The product's own crawler list

The 14 AI crawlers come from the list Verand's site audit watches on every customer site, written into this page when it's built, never typed in by hand.

Says what the file can't do

Not security, not noindex, no Crawl-delay for Google, and a list of 14 crawlers rather than every bot. The limits sit beside the builder, not in a footnote.

$0, no signup, no cap

The builder runs in your browser, so a click costs nothing and is never metered. Reading your live file is the one server call, limited to 20 a minute per visitor as a courtesy to the sites being read.

Questions

Frequently Asked Questions About the Robots.txt Generator

Where the file goes, the lines people ask about most, and what robots.txt can and can't do for you.

Where do I upload the robots.txt file, and does it work in a subfolder?

At the root of the host, so it answers at https://yourdomain.com/robots.txt. Crawlers never look for it anywhere else, so a copy in a subfolder is ignored. Each subdomain needs its own file, because the rules apply only to the host, protocol and port that serve them. On WordPress and most hosted platforms you can upload it to the site's root folder or edit it through an SEO plugin; after it's live, load the address in a browser to confirm it comes back as plain text.

What does "Disallow: /" do, and when would I ever want it?

It asks the crawlers the group names to request nothing on the site. Under User-agent: * that means every crawler without its own group, Googlebot included. It's right for a staging or development copy and almost never right on a live site. Under a single crawler's name, such as GPTBot, it keeps just that crawler out. It isn't a lock, since the file is public and honouring it is voluntary, and it doesn't remove pages from Google's index.

Should I block /wp-admin/ in robots.txt?

WordPress already does, in the virtual robots.txt it serves when you don't have a real file: Disallow: /wp-admin/ with an Allow for /wp-admin/admin-ajax.php, which the public site calls. The WordPress preset writes the same two lines. Don't add /wp-includes/ or /wp-content/, since they hold the CSS and JavaScript your pages load. The admin area is protected by its login, not by robots.txt.

Does robots.txt keep a page out of Google?

No. It controls crawling, not indexing. A blocked URL can still appear in results if other pages link to it, shown without a description. To keep a page out of the index, let Google crawl it and serve a noindex meta tag or X-Robots-Tag header; a Disallow on the same page stops Google from ever seeing that tag. The Noindex Checker shows whether a page carries one.

Is Crawl-delay worth adding?

Only if a crawler that honours it is putting real load on your server. Google says it doesn't support the field, so it never slows Googlebot. Anthropic says its crawlers support it. If you add one, the generator copies it into every Allow group it writes, because a crawler with its own group only reads the Crawl-delay in that group.

Do I need separate rules for AI crawlers like GPTBot and ClaudeBot?

Only if you want them treated differently from everyone else. With no group of their own they follow your User-agent: * rules. Give one its own group and it stops reading * entirely, which is why this generator repeats your * rules inside every Allow group it writes. To block a crawler outright, set it to Block. For a policy by purpose, training versus search versus user fetches, use the Robots.txt Generator for AI.

After the file

Write like you. Rank on Google and in AI answers from ChatGPT, Gemini, Google AI Overviews, Google AI Mode, Perplexity and Claude. Never publish a line your rules flag.

Verand writes from your own experience, gets you found on Google and in AI answers, and tracks where you rank and which answers name you. Every draft passes a compliance gate before anyone can publish it: your industry's rules if you have them, a truth-in-advertising check if you don't.

Researched from the regulators' own text and tested by Verand. Not reviewed by a licensed attorney. Your counsel confirms applicability. Not legal advice. Example shown is illustrative.

Compliance gate
Draft, paragraph 3

The 2019 fund returned 14.2% net to investors, and this offering carries guaranteed returns of 12 to 15% over a five year hold.

U.S. Securities and Exchange CommissionHard rule

Performance advertising can’t promise a return. The claim needs a basis and the required disclosures, and it still can’t be stated as a guarantee.

Ruleregd-specific-roi-no-basis
Packus-reg-d
BasisMarketing Rule 206(4)-1(d)(6); Securities Act §17(a)
ChatGPT
Gemini
AI Overviews
AI Mode
Perplexity
Claude
Recorded ChatGPT · today 09:02
Cited
Pagewillowdaleequity.com/blog/syndication-minimums
Position#2 of 3Also cited
Share of voice by engine

Willowdale Equity and four competitors · 14 days

YouCrowdStreetEquityMultipleFundriseRealtyMogul
Day 1
Across six engines31%You 22%21%15%11%
Position-weighted mentions across active prompts
The answer · a buyer's question on ChatGPT, read and recorded0:00 / 0:30
Verand

Content built to rank in Google and get cited by ChatGPTPerplexityGeminiClaude, with every claim checked before it goes live.

support@verand.ai

© 2026 Verand. All rights reserved. TermsPrivacyAI policyAccessibilitySecurity
Not legal advice. Compliance packs are researched from the regulators' own text and tested by Verand, not reviewed by a licensed attorney.