Home›Academy›Technical SEO›Robots.txt

Technical SEO · Lesson 06

Robots.txt

How robots.txt directives control crawler access by user-agent, and why blocking a URL from crawling is different from telling search engines not to index it.

06 Technical SEO12 min Read timeBeginner Level

06 / Technical SEO

A robots.txt file tells automated crawlers which parts of a website they are allowed or not allowed to request.

Robots.txt is a crawl-control mechanism. It shapes how Googlebot, Bingbot, and other crawlers spend their requests on your site, which matters for crawl budget on large websites. It does not directly remove a URL from Google's index, and confusing crawl control with index control is one of the most common technical SEO mistakes.

This lesson covers user-agents, Allow and Disallow rules, the Sitemap directive, the difference between robots.txt and noindex, and how to use a technical SEO audit with SiteAuditLint to find important URLs blocked by accident.

1. What is robots.txt?

The file follows the Robots Exclusion Protocol and must sit at the root of the host it controls:

https://www.example.com/robots.txt

A robots.txt file only applies to the exact protocol and host where it lives. The file at https://www.example.com/robots.txt does not control https://shop.example.com/ or http://www.example.com/. Each subdomain needs its own file.

A basic example looks like this:

User-agent: *
Disallow: /private/

The rule tells all crawlers represented by * not to crawl URLs under /private/.

Crawl control, not index control

Robots.txt decides whether a crawler may request a URL. It does not decide whether that URL can appear in search results. Keep that distinction in mind for every rule you write.

How a crawler uses robots.txt

FetchFile requestedBefore crawling, the crawler requests /robots.txt for that host.
MatchGroup selectedThe crawler finds the most specific User-agent group that matches it.
CrawlRules appliedEach URL is checked against Allow and Disallow paths before it is requested.

If /robots.txt returns a 4xx, Google treats the site as having no crawl restrictions. A 5xx response can cause Google to pause crawling temporarily.

2. User-agents

A user-agent identifies the crawler that a group of rules applies to:

User-agent: *
Disallow: /admin/

The * wildcard means the rule applies broadly to crawlers. A file can also contain a group for a specific crawler:

User-agent: Googlebot
Disallow: /private/

Different crawlers can therefore receive different instructions. A crawler follows only the most specific group that matches its name, so if a Googlebot group exists, Googlebot ignores the * group entirely rather than combining the two.

Common examples

User-agentPurpose
*General rule for crawlers
GooglebotGoogle's main web crawler
BingbotBing's web crawler
Googlebot-ImageGoogle's image crawler

The important point is that the user-agent determines which crawler receives the rule.

3. Allow rules

An Allow rule identifies a path a crawler may access when a broader restriction also exists:

User-agent: *
Disallow: /products/
Allow: /products/example.html

The configuration blocks the /products/ directory while explicitly allowing the specified URL. When rules conflict, Google applies the most specific rule, meaning the one with the longest matching path. If two matching rules are equally long, Google uses the less restrictive Allow rule.

Allow rules are most useful when a website needs a precise exception within a broader crawl restriction.

Practical check

When reviewing an Allow rule, ask:

  • Which user-agent does it apply to?
  • What broader Disallow rule exists?
  • Which exact URL or path is being allowed?
  • Is the allowed resource actually intended to be crawlable?

4. Disallow rules

Disallow tells a crawler not to request matching URLs:

User-agent: *
Disallow: /checkout/
Disallow: /account/
Disallow: /search/

These rules can prevent crawlers from repeatedly requesting areas that are not useful for search discovery, such as internal search results and faceted URLs that would otherwise waste crawl budget.

Paths are case-sensitive, so Disallow: /Search/ does not block /search/. Google also supports * to match any sequence of characters and $ to match the end of a URL, for example Disallow: /*?sort= or Disallow: /*.pdf$.

Blocked URLs can still appear in search

If Google discovers a blocked URL through internal or external links, it may still index the URL without crawling its content. In Google Search Console this often shows up as "Indexed, though blocked by robots.txt".

5. Robots.txt vs noindex

This distinction is one of the most important concepts in crawl control.

DirectiveWhere it livesWhat it means
Disallow: /example/robots.txtDo not crawl this path
<meta name="robots" content="noindex">Page HTMLDo not include this page in the search index
X-Robots-Tag: noindexHTTP headerSame as meta noindex, useful for PDFs and other non-HTML files

The crawler generally needs to access the page to see the noindex instruction. That creates an important difference: blocking a URL prevents crawling, while noindex controls indexing.

Example

Suppose /old-page/ contains:

<meta name="robots" content="noindex">

but robots.txt contains:

Disallow: /old-page/

The crawler is prevented from accessing the page, so the noindex directive cannot be read. The URL can stay in the index indefinitely.

Removing a page from the index

AllowKeep it crawlableMake sure no Disallow rule matches the URL.
SignalAdd noindexUse the robots meta tag or an X-Robots-Tag header.
DropPage removedAfter recrawl, the URL drops out. Only then consider blocking it.

Crawl blocking and indexing controls are not interchangeable. Use each for its own job.

6. Sitemap directive

A robots.txt file can also identify an XML sitemap:

User-agent: *
Disallow: /private/

Sitemap: https://www.example.com/sitemap.xml

The Sitemap directive gives crawlers the location of the sitemap. It must be an absolute URL, it is independent of any user-agent group, and a file can list more than one sitemap or a sitemap index. The sitemap location can also be submitted through Google Search Console and Bing Webmaster Tools. For what belongs inside the file, see Lesson 05: XML sitemaps.

Check the sitemap URL

Make sure the sitemap:

  • Loads successfully.
  • Uses the correct canonical domain.
  • Contains valid URLs.
  • Uses HTTPS when the site uses HTTPS.
  • Does not point to a blocked location.
  • Does not contain large numbers of non-indexable URLs.

7. Crawl restrictions

Robots.txt can help control crawler access to parts of a website that provide little search value:

Disallow: /admin/
Disallow: /login/
Disallow: /cart/
Disallow: /checkout/

Crawl restrictions can reduce unnecessary crawler requests, but broad blocking can also create SEO problems. A single broad rule can affect thousands of URLs. Be especially careful with rules affecting:

  • Important landing pages.
  • Product pages.
  • Category pages.
  • JavaScript and CSS resources.
  • Images.
  • Internal search pages.
  • XML sitemaps.
  • Canonical URLs.
Robots.txt is not a security tool

The file is public, and anyone can read it. Listing sensitive paths in robots.txt advertises them. Protect private areas with authentication instead.

8. Common robots.txt problems

Blocking important pages

Disallow: /products/ prevents crawlers from accessing primary commercial pages if that directory holds them.

Blocking the entire website

Disallow: / under User-agent: * blocks crawling across the whole site. This often ships by mistake when a staging robots.txt goes live.

Blocking resources

Rules that stop access to important CSS, JavaScript, or image files can prevent Google from rendering pages correctly.

Sitemap pointing to the wrong location

Sitemap: https://example.com/sitemap-old.xml gives an incorrect discovery path if that file no longer exists.

Using robots.txt to deindex

Blocking a URL that carries noindex hides the directive from crawlers, so the page can remain indexed.

9. How to audit robots.txt with SiteAuditLint

A website crawler such as SiteAuditLint reveals URLs affected by robots.txt restrictions and helps you check whether important pages are being blocked.

SiteAuditLint robots.txt workflow

  1. Read the fileOpen /robots.txt on every host and note each user-agent group, rule, and Sitemap line.
  2. Crawl the siteRun a crawl with SiteAuditLint so blocked URLs are recorded rather than silently skipped.
  3. Review blocked URLsLook for important pages, sitemap URLs, and CSS or JavaScript resources affected by a restriction.
  4. Group by ruleIdentify large groups of URLs blocked by the same rule, since one broad pattern often explains most issues.
  5. Fix the rulesNarrow overly broad paths, add Allow exceptions, or switch to noindex where indexing is the real goal.
  6. RecheckRecrawl, then confirm the live file in the Google Search Console robots.txt report.
Intentional beats minimal

The goal is not to remove every robots.txt restriction. It is to make sure the restrictions match the site's intended crawl strategy.

10. Practical example

Imagine an online store with /products/, /cart/, and /checkout/. Its robots.txt contains:

User-agent: *
Disallow: /cart/
Disallow: /checkout/

Sitemap: https://www.example.com/sitemap.xml

Compare that with a single-line mistake:

PathIntended setupDisallow: /
/products/CrawlableBlocked
/cart/BlockedBlocked
/checkout/BlockedBlocked
SitemapDeclared and reachableNot declared

The first configuration keeps crawlers out of cart and checkout URLs, keeps product pages available for crawling, and points to the XML sitemap. The second blocks the entire site.

11. Exercise

Open the robots.txt file for a website you manage and answer these questions:

  1. User-agentsWhich user-agents are defined?
  2. Disallow rulesAre there any Disallow rules?
  3. Allow rulesAre there any Allow rules?
  4. SitemapIs an XML sitemap specified?
  5. Important sectionsAre important sections of the site blocked?
  6. Sitemap URLsAre sitemap URLs affected by a crawl restriction?
  7. IndexingAre you using robots.txt to solve an indexing problem that should use noindex or another indexing signal?

Record any rules that could affect important URLs.

12. Robots.txt checklist

Before finishing a technical SEO audit, work through this list.

Robots.txt audit checklist

0 of 10 tasks completed

13. Key takeaways

  • Robots.txt controls crawling, not indexing. A Disallow rule does not guarantee a URL disappears from search results.
  • User-agent determines who receives the rule. Rules can apply broadly or target a specific crawler, and a crawler follows only its most specific group.
  • Allow and Disallow rules work together. The longest matching path wins, which lets you create precise exceptions.
  • Sitemaps support discovery. The Sitemap directive tells crawlers where an XML sitemap can be found.
  • Blocked does not mean noindexed. A crawler must be able to access a page to read its noindex directive.

Knowledge check

1. What does Disallow: /private/ do?
2. A page has noindex and is also blocked in robots.txt. What happens?
3. A file has Disallow: /products/ and Allow: /products/example.html. Can Googlebot crawl /products/example.html?
4. Which rule blocks the entire website for all crawlers?

Ready to go further? Review the technical SEO checklist, or revisit Lesson 05: XML sitemaps.