02 / Crawl control
This page defines 29 Crawling & Indexing terms. Each entry gives a plain definition, a What to check step you can apply to your own site, and links to deeper SiteAuditLint guides where they exist.
Status codes
HTTP status codes are three-digit responses a server returns for each request, telling browsers and crawlers whether a URL succeeded (2xx), redirected (3xx), was not found (4xx), or failed (5xx).
Important pages return 200, internal links do not point to 3xx or 4xx URLs, and 5xx errors are not spiking.
Go deeper: Status codes, HTTP status codes and SEO
301 redirect
A 301 redirect permanently sends users and search engines from one URL to another, and passes ranking signals to the destination.
Each old URL redirects in a single hop to the most relevant new URL, not to the homepage by default.
Does a 301 redirect lose link equity?
Google has stated that 3xx redirects do not lose PageRank, though redirecting to irrelevant pages can still weaken results.
Go deeper: Redirects
302 redirect
A 302 redirect temporarily sends users to another URL while signaling that the original should remain the indexed version.
Use 302s only for genuinely temporary moves, and switch to 301 if the change becomes permanent.
301 vs 302 redirect?
A 301 says the move is permanent and consolidates signals to the new URL. A 302 says it is temporary, so Google may keep the original URL indexed.
Go deeper: How to find and fix redirect problems
404 error
A 404 error is the status code returned when a requested page does not exist. 404s are normal and do not harm a site's rankings on their own.
Fix internal links pointing to 404s, and redirect removed pages that still have backlinks or traffic.
Go deeper: Broken link audit
Soft 404
A soft 404 is a page that tells users it does not exist, or has almost no content, but returns a 200 status code instead of 404.
Review soft 404s in the Page indexing report and return a true 404 or 410, or add real content.
Robots.txt
Robots.txt is a text file at the root of a domain that tells crawlers which paths they may or may not crawl. It controls crawling, not indexing.
No important sections or resource files are disallowed, and the file returns a 200 status.
Does robots.txt stop a page from being indexed?
No. A blocked URL can still be indexed if other pages link to it. Use noindex, and leave the page crawlable so Google can see it.
Go deeper: Robots.txt, Robots.txt checklist
Source: Google Search Central: Introduction to robots.txt
XML sitemap
An XML sitemap is a file listing the URLs you want search engines to discover and crawl, optionally with last-modified dates.
The sitemap only lists indexable, canonical URLs that return 200, and it is submitted in Search Console.
Go deeper: XML sitemaps, XML sitemap errors
Source: Google Search Central: Sitemaps overview
Meta robots tag
A meta robots tag is an HTML element in the page head that gives search engines page-level directives, such as noindex or nofollow.
Templates do not ship with leftover noindex tags from staging or development environments.
Source: Google Search Central: Robots meta tag specifications
Noindex
Noindex is a directive telling search engines not to include a page in their index. It can be set in a meta robots tag or an X-Robots-Tag header.
Noindexed pages are not blocked in robots.txt, are not in your XML sitemap, and are not pages you want to rank.
Go deeper: Noindex issues: how to find and fix them
Indexability
Indexability is whether a page is eligible to be indexed. A page is indexable when it returns 200, is crawlable, has no noindex directive, and is not canonicalized to another URL.
Run an indexability check on key templates after every release, since a single change can remove pages from search.
Go deeper: Indexability checklist
Canonical tag
A canonical tag is an HTML link element that tells search engines which URL is the preferred version when several URLs show the same or very similar content. It is a strong hint, not a directive.
The canonical points to the intended URL, returns 200, uses the correct protocol and hostname, and does not conflict with redirects, noindex, or hreflang.
Canonical tag vs 301 redirect?
A redirect sends users and crawlers to another URL. A canonical keeps both URLs accessible but consolidates ranking signals to one.
Why is Google ignoring my canonical tag?
Google may choose a different canonical when other signals disagree, such as internal links, sitemaps, or redirects pointing to a different version.
Go deeper: Canonicals, Canonical tag issues and fixes
Source: Google Search Central: Consolidate duplicate URLs
Self-referencing canonical
A self-referencing canonical is a canonical tag that points to the page's own URL, confirming it as the preferred version.
Parameterized and paginated URLs do not accidentally self-canonicalize when they should point to a main version.
Go deeper: Canonical testing
Crawl depth
Crawl depth is the number of clicks needed to reach a page from the homepage. Deeply buried pages tend to be crawled less often.
Keep important pages within about three clicks of the homepage through navigation and internal links.
Go deeper: Deep page issue
Orphan page
An orphan page is a page with no internal links pointing to it, making it hard for users and crawlers to find.
Compare sitemap and analytics URLs with crawled URLs to find pages that no internal link reaches.
Go deeper: Orphan pages: how to find and resolve them
Redirect chain
A redirect chain occurs when a URL redirects to another URL that redirects again before reaching the final destination.
Update redirects and internal links so every URL reaches its destination in a single hop.
Go deeper: Redirect chain issue
Redirect loop
A redirect loop occurs when URLs redirect to each other in a cycle, so the page never loads.
Test redirect rules after server or CMS changes, especially HTTP to HTTPS and trailing slash rules.
Go deeper: Redirect loop issue
Hreflang
Hreflang is an attribute that tells search engines which language or regional version of a page to show to which users.
Language and region codes are valid, every version references the others, and each hreflang URL is canonical and indexable.
Go deeper: International SEO checklist
Hreflang return tags
Hreflang return tags are the reciprocal references required between alternate pages. If page A lists page B, page B must list page A, or the annotation may be ignored.
Audit hreflang clusters for missing return links after adding or removing language versions.
Googlebot
Googlebot is Google's web crawler, with smartphone and desktop variants. Google crawls most sites primarily with the smartphone crawler.
Verify Googlebot traffic through reverse DNS lookups, since fake bots often spoof its user agent.
Source: Google Search Central: Googlebot
Canonicalization
Canonicalization is the process by which a search engine chooses one representative URL from a group of duplicate or near-duplicate pages.
Compare your declared canonical with the Google-selected canonical in the URL Inspection tool.
Source: Google Search Central: Consolidate duplicate URLs
Duplicate URL clustering
Duplicate URL clustering is how search engines group URLs with the same content before choosing a canonical, for example versions with tracking parameters or different capitalization.
Normalize URL variants with redirects and consistent internal linking so clusters stay small.
X-Robots-Tag
The X-Robots-Tag is an HTTP header that carries robots directives, such as noindex, for any file type, including PDFs and images.
Inspect response headers on non-HTML files and confirm no unintended noindex is set at the server level.
Go deeper: HTTP header SEO testing
Source: Google Search Central: Robots meta tag and X-Robots-Tag
Crawl budget
Crawl budget is the number of URLs a search engine can and wants to crawl on a site within a given time. It mainly matters for large sites or sites that generate many URLs.
Reduce crawl waste from parameters, faceted URLs, and redirect chains, and monitor Crawl Stats for slow response times.
Does crawl budget matter for small sites?
Usually not. Google says most sites with fewer than a few thousand URLs are crawled efficiently.
Source: Google Search Central: Crawl budget management for large sites
Crawl trap
A crawl trap is a site structure that generates a near-infinite number of URLs, such as endless calendar pages or combinable filters, wasting crawl resources.
Look for URL patterns in crawl data and server logs that grow without limit.
Index bloat
Index bloat happens when a site has many low-value pages indexed, such as filter pages, tag archives, or internal search results.
Compare the number of indexed pages with the number of pages you actually want to rank.
Faceted navigation
Faceted navigation lets users filter listings by attributes such as size, color, or price, often creating many URL combinations.
Decide which facet combinations have search demand, and keep the rest out of crawling and indexing.
Go deeper: Ecommerce SEO checklist
Parameter handling
Parameter handling is how a site manages URL parameters for tracking, sorting, and filtering so they do not create duplicate or wasteful URLs.
Internal links use clean URLs, and parameter versions canonicalize to the main version.
Go deeper: URL parameters issue
Pagination
Pagination splits content across numbered pages, such as category pages 1, 2, and 3. Google no longer uses rel=next and rel=prev as indexing signals.
Paginated pages are crawlable through normal links and usually self-canonicalize rather than pointing to page 1.
Site migration
A site migration is any major change to a site's URLs, domain, platform, or structure that can affect search visibility.
Build a full redirect map, benchmark rankings beforehand, and run post-launch checks on redirects, canonicals, and indexability.
Go deeper: Website migration SEO checklist, Redirect map validation
Sources
Official documentation referenced on this page:
- Google Search Central: Introduction to robots.txt
- Google Search Central: Sitemaps overview
- Google Search Central: Robots meta tag and X-Robots-Tag
- Google Search Central: Consolidate duplicate URLs
- Google Search Central: Googlebot
- Google Search Central: Crawl budget management for large sites
Browse the full SEO terms glossary, or run a technical SEO audit with SiteAuditLint to check these signals on your own site.