- Robots.txt manages crawling. It does not remove pages from the index.
- The file must live at the root of each host and protocol, and each subdomain needs its own.
- The most specific matching rule wins, and a crawler obeys only the most specific user-agent group that matches it.
- Do not block CSS or JavaScript that is needed to render the page.
- A 5xx response on robots.txt can pause crawling of the whole site.
Robots.txt is a text file that sits at the root of a website and tells crawlers which parts of the site they are allowed to request. It follows the Robots Exclusion Protocol, formalised as RFC 9309 in 2022. Reputable crawlers obey it. Bad ones ignore it, so never treat it as a way to hide anything sensitive.
Where it lives
The file must be at the root of the host, for example https://www.example.com/robots.txt. It applies only to that exact protocol, hostname and port, so shop.example.com needs its own file.
Syntax
User-agent: *
Disallow: /basket/
Disallow: /search
Allow: /search/help
User-agent: GPTBot
Disallow: /
Sitemap: https://www.example.com/sitemap.xml
- User-agent names the crawler a group of rules applies to. An asterisk means all crawlers.
- Disallow blocks any URL path that starts with the value given.
- Allow creates an exception within a disallowed path.
- Sitemap gives the absolute URL of an XML sitemap and can be listed more than once.
- * matches any sequence of characters and $ marks the end of a URL, so Disallow: /*.pdf$ blocks URLs ending in .pdf.
How rules are resolved
- A crawler follows only one group: the one with the most specific matching user-agent. If you add a Googlebot group, Googlebot ignores the rules under the asterisk.
- Within a group, the longest matching path wins. When an Allow and a Disallow are equally specific, Google uses the less restrictive one.
- Paths are case-sensitive. /Folder/ and /folder/ are different.
What robots.txt does not do
- It does not deindex. A disallowed URL can still appear in results, without a description, if other pages link to it. To remove a page, allow crawling and use noindex.
- It does not support noindex. Google stopped honouring noindex lines in robots.txt in 2019.
- It does not pass or block link equity in any tidy way. Links on a blocked page simply are not seen.
- Crawl-delay is ignored by Google, although Bing and some others respect it.
How status codes are handled
- 2xx: the file is read and followed. Google caches it for up to 24 hours.
- 4xx: treated as if no file exists, so everything may be crawled.
- 5xx: Google treats the site as temporarily off limits and stops crawling. If the errors continue for a long period it may fall back to a cached copy or eventually assume there are no restrictions.
Google reads only the first 500 KiB of the file. If yours is anywhere near that, the rules need simplifying.
Common mistakes
- A leftover Disallow: / from staging going live.
- Blocking JavaScript, CSS or API endpoints that the page needs to render.
- Blocking a URL and adding noindex to it, so the noindex is never seen.
- Blocking parameter URLs that carry canonicals, so the canonical cannot be read.
- Trying to keep private files out of sight. The file is public and lists exactly where you do not want people to look.
- Forgetting the trailing slash, so Disallow: /blog also blocks /blog-archive.
What to block
Good candidates are internal search results, endless filter and sort combinations, basket and checkout steps, and other spaces that generate near-infinite URLs with no search value. On small sites you often need very little in there.
How to test
- The robots.txt report in Search Console shows the versions Google fetched and any parsing errors.
- URL Inspection tells you whether a specific URL is blocked.
- Crawling tools let you test a draft file against a list of URLs before it goes live.
The GEO angle
Robots.txt is now the main place where you decide which AI crawlers can access your content. Training crawlers, search crawlers and user-triggered fetchers have different user agents, so you can allow one and block another. If you want to appear in AI answers, check that your file, your CDN and your firewall are not quietly blocking the bots that power them.
Add your title here
This is a paragraph. Writing in paragraphs lets visitors find what they are looking for quickly and easily. Make sure the title suits the content of this text.




