06 / Technical SEO
A robots.txt file tells automated crawlers which parts of a website they are allowed or not allowed to request.
Robots.txt is a crawl-control mechanism. It shapes how Googlebot, Bingbot, and other crawlers spend their requests on your site, which matters for crawl budget on large websites. It does not directly remove a URL from Google's index, and confusing crawl control with index control is one of the most common technical SEO mistakes.
This lesson covers user-agents, Allow and Disallow rules, the Sitemap directive, the difference between robots.txt and noindex, and how to use a technical SEO audit with SiteAuditLint to find important URLs blocked by accident.
1. What is robots.txt?
The file follows the Robots Exclusion Protocol and must sit at the root of the host it controls:
https://www.example.com/robots.txt
A robots.txt file only applies to the exact protocol and host where it lives. The file at https://www.example.com/robots.txt does not control https://shop.example.com/ or http://www.example.com/. Each subdomain needs its own file.
A basic example looks like this:
User-agent: *
Disallow: /private/
The rule tells all crawlers represented by * not to crawl URLs under /private/.
Robots.txt decides whether a crawler may request a URL. It does not decide whether that URL can appear in search results. Keep that distinction in mind for every rule you write.
How a crawler uses robots.txt
/robots.txt for that host.User-agent group that matches it.Allow and Disallow paths before it is requested.If /robots.txt returns a 4xx, Google treats the site as having no crawl restrictions. A 5xx response can cause Google to pause crawling temporarily.
2. User-agents
A user-agent identifies the crawler that a group of rules applies to:
User-agent: *
Disallow: /admin/
The * wildcard means the rule applies broadly to crawlers. A file can also contain a group for a specific crawler:
User-agent: Googlebot
Disallow: /private/
Different crawlers can therefore receive different instructions. A crawler follows only the most specific group that matches its name, so if a Googlebot group exists, Googlebot ignores the * group entirely rather than combining the two.
Common examples
| User-agent | Purpose |
|---|---|
* | General rule for crawlers |
Googlebot | Google's main web crawler |
Bingbot | Bing's web crawler |
Googlebot-Image | Google's image crawler |
The important point is that the user-agent determines which crawler receives the rule.
3. Allow rules
An Allow rule identifies a path a crawler may access when a broader restriction also exists:
User-agent: *
Disallow: /products/
Allow: /products/example.html
The configuration blocks the /products/ directory while explicitly allowing the specified URL. When rules conflict, Google applies the most specific rule, meaning the one with the longest matching path. If two matching rules are equally long, Google uses the less restrictive Allow rule.
Allow rules are most useful when a website needs a precise exception within a broader crawl restriction.
Practical check
When reviewing an Allow rule, ask:
- Which user-agent does it apply to?
- What broader
Disallowrule exists? - Which exact URL or path is being allowed?
- Is the allowed resource actually intended to be crawlable?
4. Disallow rules
Disallow tells a crawler not to request matching URLs:
User-agent: *
Disallow: /checkout/
Disallow: /account/
Disallow: /search/
These rules can prevent crawlers from repeatedly requesting areas that are not useful for search discovery, such as internal search results and faceted URLs that would otherwise waste crawl budget.
Paths are case-sensitive, so Disallow: /Search/ does not block /search/. Google also supports * to match any sequence of characters and $ to match the end of a URL, for example Disallow: /*?sort= or Disallow: /*.pdf$.
If Google discovers a blocked URL through internal or external links, it may still index the URL without crawling its content. In Google Search Console this often shows up as "Indexed, though blocked by robots.txt".
5. Robots.txt vs noindex
This distinction is one of the most important concepts in crawl control.
| Directive | Where it lives | What it means |
|---|---|---|
Disallow: /example/ | robots.txt | Do not crawl this path |
<meta name="robots" content="noindex"> | Page HTML | Do not include this page in the search index |
X-Robots-Tag: noindex | HTTP header | Same as meta noindex, useful for PDFs and other non-HTML files |
The crawler generally needs to access the page to see the noindex instruction. That creates an important difference: blocking a URL prevents crawling, while noindex controls indexing.
Example
Suppose /old-page/ contains:
<meta name="robots" content="noindex">
but robots.txt contains:
Disallow: /old-page/
The crawler is prevented from accessing the page, so the noindex directive cannot be read. The URL can stay in the index indefinitely.
Removing a page from the index
Disallow rule matches the URL.X-Robots-Tag header.Crawl blocking and indexing controls are not interchangeable. Use each for its own job.
6. Sitemap directive
A robots.txt file can also identify an XML sitemap:
User-agent: *
Disallow: /private/
Sitemap: https://www.example.com/sitemap.xml
The Sitemap directive gives crawlers the location of the sitemap. It must be an absolute URL, it is independent of any user-agent group, and a file can list more than one sitemap or a sitemap index. The sitemap location can also be submitted through Google Search Console and Bing Webmaster Tools. For what belongs inside the file, see Lesson 05: XML sitemaps.
Check the sitemap URL
Make sure the sitemap:
- Loads successfully.
- Uses the correct canonical domain.
- Contains valid URLs.
- Uses HTTPS when the site uses HTTPS.
- Does not point to a blocked location.
- Does not contain large numbers of non-indexable URLs.
7. Crawl restrictions
Robots.txt can help control crawler access to parts of a website that provide little search value:
Disallow: /admin/
Disallow: /login/
Disallow: /cart/
Disallow: /checkout/
Crawl restrictions can reduce unnecessary crawler requests, but broad blocking can also create SEO problems. A single broad rule can affect thousands of URLs. Be especially careful with rules affecting:
- Important landing pages.
- Product pages.
- Category pages.
- JavaScript and CSS resources.
- Images.
- Internal search pages.
- XML sitemaps.
- Canonical URLs.
The file is public, and anyone can read it. Listing sensitive paths in robots.txt advertises them. Protect private areas with authentication instead.
8. Common robots.txt problems
Blocking important pages
Disallow: /products/ prevents crawlers from accessing primary commercial pages if that directory holds them.
Blocking the entire website
Disallow: / under User-agent: * blocks crawling across the whole site. This often ships by mistake when a staging robots.txt goes live.
Blocking resources
Rules that stop access to important CSS, JavaScript, or image files can prevent Google from rendering pages correctly.
Sitemap pointing to the wrong location
Sitemap: https://example.com/sitemap-old.xml gives an incorrect discovery path if that file no longer exists.
Using robots.txt to deindex
Blocking a URL that carries noindex hides the directive from crawlers, so the page can remain indexed.
9. How to audit robots.txt with SiteAuditLint
A website crawler such as SiteAuditLint reveals URLs affected by robots.txt restrictions and helps you check whether important pages are being blocked.
SiteAuditLint robots.txt workflow
- Read the fileOpen
/robots.txton every host and note each user-agent group, rule, andSitemapline. - Crawl the siteRun a crawl with SiteAuditLint so blocked URLs are recorded rather than silently skipped.
- Review blocked URLsLook for important pages, sitemap URLs, and CSS or JavaScript resources affected by a restriction.
- Group by ruleIdentify large groups of URLs blocked by the same rule, since one broad pattern often explains most issues.
- Fix the rulesNarrow overly broad paths, add
Allowexceptions, or switch tonoindexwhere indexing is the real goal. - RecheckRecrawl, then confirm the live file in the Google Search Console robots.txt report.
The goal is not to remove every robots.txt restriction. It is to make sure the restrictions match the site's intended crawl strategy.
10. Practical example
Imagine an online store with /products/, /cart/, and /checkout/. Its robots.txt contains:
User-agent: *
Disallow: /cart/
Disallow: /checkout/
Sitemap: https://www.example.com/sitemap.xml
Compare that with a single-line mistake:
| Path | Intended setup | Disallow: / |
|---|---|---|
/products/ | Crawlable | Blocked |
/cart/ | Blocked | Blocked |
/checkout/ | Blocked | Blocked |
| Sitemap | Declared and reachable | Not declared |
The first configuration keeps crawlers out of cart and checkout URLs, keeps product pages available for crawling, and points to the XML sitemap. The second blocks the entire site.
11. Exercise
Open the robots.txt file for a website you manage and answer these questions:
- User-agentsWhich user-agents are defined?
- Disallow rulesAre there any
Disallowrules? - Allow rulesAre there any
Allowrules? - SitemapIs an XML sitemap specified?
- Important sectionsAre important sections of the site blocked?
- Sitemap URLsAre sitemap URLs affected by a crawl restriction?
- IndexingAre you using robots.txt to solve an indexing problem that should use
noindexor another indexing signal?
Record any rules that could affect important URLs.
12. Robots.txt checklist
Before finishing a technical SEO audit, work through this list.
Robots.txt audit checklist
0 of 10 tasks completed
13. Key takeaways
- Robots.txt controls crawling, not indexing. A
Disallowrule does not guarantee a URL disappears from search results. - User-agent determines who receives the rule. Rules can apply broadly or target a specific crawler, and a crawler follows only its most specific group.
- Allow and Disallow rules work together. The longest matching path wins, which lets you create precise exceptions.
- Sitemaps support discovery. The
Sitemapdirective tells crawlers where an XML sitemap can be found. - Blocked does not mean noindexed. A crawler must be able to access a page to read its
noindexdirective.
Knowledge check
Ready to go further? Review the technical SEO checklist, or revisit Lesson 05: XML sitemaps.