A robots.txt file is a plain-text file at the root of your site that tells crawlers which URLs they may request. It controls crawling, not indexing. Google’s own robots.txt introduction says it is “used mainly to avoid overloading your site with requests” and that “it is not a mechanism for keeping a web page out of Google.” If you want a page out of search results, you need noindex, a password or deletion, not a Disallow line.
That one distinction causes most of the robots.txt mistakes I find. When I audit a site that “lost” pages, a Disallow meant to hide them is often the reason they’re still showing. Below I cover what the file does, the syntax Google actually follows, which rule wins when two match, how to test it, AI crawlers, and the WordPress and Shopify versions, with example files you can copy.
What Does Robots.txt Do, and What Doesn’t It Do?
Robots.txt asks crawlers not to fetch certain URLs. Well-behaved bots like Googlebot respect it. That’s the whole job.
What it doesn’t do is remove anything. Google’s intro page says a disallowed page “can still be indexed if linked to from other sites,” and the URL plus anchor text from those links can still appear in results. So a blocked page can rank as a bare URL with no description, which is often worse than either option you wanted.
It also isn’t security. Google notes that not every search engine supports robots.txt rules, and different crawlers interpret the syntax differently. Anyone can open your robots.txt in a browser, so listing /secret-admin-panel/ in it is basically a signpost.
What Decides How Your Robots.txt Should Look?
Four variables change the right file for your site. I check them before I write a single line.
- Your goal. Reducing crawl waste, hiding a page and controlling AI bots need three different tools. Only the first and third are robots.txt jobs.
- Your platform. WordPress serves a virtual file, Shopify uses a Liquid template, and custom builds need a real file on the server.
- Your hosts. Each subdomain, protocol and port needs its own file.
- Site size. On small sites, the default file is usually right. On large stores, filter and search URLs make custom rules worth the effort.
If your goal is keeping pages out of Google, then use noindex and leave the page crawlable. If your goal is stopping Googlebot from wasting requests on endless filter or search URLs, then robots.txt is the right tool. If you’re not sure, then leave the default file alone.
How Is a Robots.txt File Structured?
A file is made of groups. Each group starts with one or more User-agent lines, followed by Disallow and Allow rules. Google’s robots.txt specification supports four fields: user-agent, allow, disallow and sitemap. It says other fields such as crawl-delay aren’t supported.
Here’s a correct basic file for a typical small site. The * group applies to every crawler that doesn’t have a group of its own.
User-agent: *
Disallow: /search/
Disallow: /cart/
Allow: /
Sitemap: https://www.example.com/sitemap.xml
A few rules apply everywhere. Paths must start with / and are case-sensitive, so /Cart/ and /cart/ are different. User-agent names aren’t case-sensitive. The Sitemap line takes a full URL and can sit anywhere in the file.
One old trick no longer works. Google retired support for noindex inside robots.txt on September 1, 2019, according to its note on unsupported rules. If you still see Noindex: lines in a file, they do nothing for Google.
Wildcards: * and $
Google supports two special characters. * means zero or more of any character, and $ marks the end of the URL. Rules match from the start of the path, query string included.
User-agent: *
Disallow: /*?sort=
Disallow: /*.pdf$
Disallow: /fish/
In order, those three rules block any URL containing a sort parameter, any URL that ends in .pdf, and the /fish/ folder with everything in it.
The spec gives useful examples. /*.php$ matches /filename.php but not /filename.php?parameters, because the URL no longer ends in .php. And /fish matches /fishheads too, which surprises people. Add the trailing slash when you mean a folder.
Which Rule Wins When Two Rules Match?
Google’s precedence rule is simple once you see it. The most specific rule wins, measured by the length of the rule path. When two matching rules conflict and are equally specific, Google uses the least restrictive one, which means Allow.
| URL | Rules | Rule Google applies |
|---|---|---|
/page | Allow: /p and Disallow: / | Allow: /p, it’s longer |
/folder/page | Allow: /folder and Disallow: /folder | Allow: /folder, tie goes to least restrictive |
/page.htm | Allow: /page and Disallow: /*.htm | Disallow: /*.htm, it’s longer |
/ | Allow: /$ and Disallow: / | Allow: /$, it’s more specific |
All four rows come from Google’s spec. Note that order in the file doesn’t matter for Google. Writing Allow first doesn’t make it win.
How Does Googlebot Choose Its Group?
A crawler follows only the group with the most specific user agent that matches it. So if you add a User-agent: Googlebot group, Googlebot ignores everything under User-agent: *. I’ve seen sites add a Googlebot group with one rule and accidentally unblock their whole admin area for Google.
User-agent: *
Disallow: /wp-admin/
Disallow: /search/
User-agent: Googlebot
Disallow: /wp-admin/
Disallow: /search/
Disallow: /print/
In that file, Googlebot reads only its own group, which is why the first two rules are repeated there.
Where Does the File Go, and How Big Can It Be?
The file must sit in the top-level directory, at /robots.txt. A file at /folder/robots.txt isn’t valid, because crawlers don’t look for one in subdirectories. Rules apply only to the exact host, protocol and port that serves the file, so blog.example.com and www.example.com each need their own.
Size has a hard limit. Google’s spec says it enforces a 500 KiB limit, and content after that is ignored. In my experience, a small business site rarely needs more than a few dozen lines. If your file is approaching that size, something is generating rules automatically and needs a look.
Two more details from the spec matter during incidents. Google generally caches the file for up to 24 hours, so changes aren’t instant. And status codes change behavior: a 4xx (except 429) is treated as if no robots.txt exists, while a 5xx makes Google stop crawling the site for the first 12 hours, then fall back to the last good copy for up to 30 days while it keeps trying.
How Do You Test a Robots.txt File?
The old robots.txt Tester in Search Console is gone. Google retired it in December 2023 and replaced it with the robots.txt report, which lives under Settings. Google’s robots.txt report help page says it shows the robots.txt files found for the top 20 hosts on your site, when each was last crawled, and any warnings or errors.
The report works on Domain properties and on URL-prefix properties without a path. It also has a “Request a recrawl” option for emergencies, like after you’ve removed a rule that blocked your whole site.
For single URLs, I use URL Inspection. It tells you whether crawling is allowed for that exact URL. Before I ship a new rule, I also list ten real URLs it should block and ten it shouldn’t, then check each one by hand against the precedence rules above. That takes five minutes and has saved me from more than one bad wildcard.
How Should You Handle AI Crawlers in Robots.txt?
Each AI company runs its own crawlers with their own tokens, and blocking one doesn’t block the others. The tokens I use in robots.txt are GPTBot and OAI-SearchBot (OpenAI), ClaudeBot and Claude-SearchBot (Anthropic), PerplexityBot (Perplexity) and Google-Extended (Google’s control for Gemini training and grounding).
The usual decision is “visible in AI search, not used for training.” That file looks like this:
User-agent: OAI-SearchBot
Allow: /
User-agent: Claude-SearchBot
Allow: /
User-agent: PerplexityBot
Allow: /
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Google-Extended
Disallow: /
Don’t block Googlebot to avoid AI Overviews. You’d leave Google Search completely. The full reasoning per vendor, plus CDN settings that block bots before robots.txt is even read, is in my guide to technical SEO for AI search.
What Are the Most Common Robots.txt Mistakes?
These are the ones I see in real audits, roughly in order of damage.
Leaving Disallow: / from staging at launch. A staging site blocks everything, the file gets copied to production, and the new site vanishes from crawling. It’s one of the three staging leftovers I check first after any launch, and my website redesign without losing SEO guide covers the rest.
Blocking pages you want deindexed. Google’s noindex documentation says that for noindex to work, the page “must not be blocked by a robots.txt file.” Block it and Google never sees the tag, so the URL can stay indexed. Remove the Disallow, let Google recrawl, and only block afterwards if you still need to.
Blocking CSS and JavaScript. Google renders pages, so it needs the files that build them. Google’s intro allows blocking unimportant resources only when pages “won’t be significantly affected by the loss.” Blocking your theme folder fails that test.
Forgetting a group doesn’t inherit. A new User-agent: Googlebot group silently drops every rule under * for Googlebot, as shown above.
Mixing up case or folders. Disallow: /Blog doesn’t block /blog/, and Disallow: /blog also blocks /blog-tips/.
How Does Robots.txt Work on WordPress?
WordPress serves a virtual robots.txt when no physical file exists. In current WordPress core code, the default output blocks the admin folder, allows the AJAX endpoint, and adds a sitemap line when the site is public and core sitemaps are on:
User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php
Sitemap: https://www.example.com/wp-sitemap.xml
SEO plugins usually swap that sitemap line for their own. If you upload a real robots.txt file to the site root, the web server serves it and the virtual one stops showing. Most SEO plugins also include an editor for the file.
One detail many guides still get wrong. In current WordPress code, the “Discourage search engines from indexing this site” box under Settings, Reading doesn’t add Disallow: / to robots.txt. It adds a noindex robots meta tag to pages instead, which is the right tool for the job. Untick it at launch all the same.
How Does Robots.txt Work on Shopify?
Shopify generates a default robots.txt for every store, and its default rules already block things like cart, checkout, sort and multi-filter URLs. To change it, you add a robots.txt.liquid template in your theme’s Templates folder. Shopify’s robots.txt developer docs strongly recommend keeping the provided Liquid objects, because the default rules are updated regularly.
This is the pattern from Shopify’s docs, adding one rule to the default * group while keeping everything else:
{% for group in robots.default_groups %}
{{- group.user_agent }}
{%- for rule in group.rules -%}
{{ rule }}
{%- endfor -%}
{%- if group.user_agent.value == '*' -%}
{{ 'Disallow: /*?q=*' }}
{%- endif -%}
{%- if group.sitemap != blank -%}
{{ group.sitemap }}
{%- endif -%}
{% endfor %}
My advice for stores is the same one I give in my Shopify posts. Leave the default alone unless Search Console shows a specific crawl problem, like a filter app generating thousands of parameter URLs.
Robots.txt Decision Matrix
| If you want to… | Then use | Not |
|---|---|---|
| Remove a page from Google | noindex, password or deletion | robots.txt Disallow |
| Stop crawl waste on filter or search URLs | robots.txt Disallow with wildcards | noindex alone |
| Stay in AI search but opt out of training | Allow search bots, disallow training bots | Blocking Googlebot |
| Hide a staging site | Password protection | robots.txt only |
| Point crawlers to your sitemap | Sitemap: line plus Search Console | Nothing else needed |
When Should You Get Help With Robots.txt?
Most sites never need more than the default file. Bring in someone who knows technical SEO when you’re writing wildcard rules for a large store, when a robots.txt change coincides with a traffic drop, or when bot rules sit in a CDN as well as the file. On large sites, robots.txt is also your main lever for crawl budget, so mistakes cost more.
That’s work my team handles in our technical SEO service. If robots.txt is one item on a longer list, my guide on how to do a technical SEO audit shows where it fits, and technical SEO for beginners explains the basics behind it.
Frequently Asked Questions
Does Every Website Need a Robots.txt File?
No. If the file is missing and the server returns a 404, Google treats the site as having no crawl restrictions. A file is still useful for a Sitemap line and for blocking crawl traps like internal search results, and most platforms create one for you anyway.
Can Robots.txt Remove a Page From Google?
No. It only stops crawling. Google says a disallowed URL can still be indexed if other sites link to it. To remove a page, allow crawling and add noindex, put it behind a password, or delete it.
How Long Does Google Take to See Robots.txt Changes?
Google generally caches robots.txt for up to 24 hours. If you’ve fixed something urgent, like an accidental sitewide block, use “Request a recrawl” in the robots.txt report in Search Console.
Does Google Support Crawl-Delay?
No. Google’s specification lists only user-agent, allow, disallow and sitemap, and says fields like crawl-delay aren’t supported. Some other search engines read it, but Googlebot adjusts its crawl rate based on how your server responds.
Last updated: October 2026 by Mizanur Rahman



