SEO Fundamentals · Lesson 02

Crawling

How crawlers discover URLs, follow links, and request pages from your website, plus how to find the technical issues that interrupt discovery.

02 Discovery9 min Read timeBeginner Level

02 / Discovery

Crawling is the process search engines use to discover and access web pages. Before a page can be considered for indexing, a search engine generally needs to find its URL and retrieve its content.

This lesson covers how crawlers find URLs, follow internal and external links, use XML sitemaps, and request pages from a website. It also explains how to inspect crawl activity and use SiteAuditLint to identify technical issues that affect website crawling.

What is crawling in SEO?

Crawling is the process of automatically discovering URLs and requesting their content using software called a crawler, bot, or spider. Search engines use crawlers to explore the web, discover new pages, revisit existing pages, and collect information that may be used during indexing. For a deeper look at the software itself, read what a website crawler is and how it works.

For example, when a search engine discovers a link from your homepage to a blog post, it can follow that link and request the blog post's URL. If the page is accessible, the crawler can retrieve its HTML and identify additional links to explore.

How crawling works

01Known URLA URL discovered through a link or sitemap
02Crawler requests pageThe crawler sends an HTTP request
03Page is retrievedThe server returns a response and page content
04Links are discoveredThe crawler can find more URLs to visit

Simplified crawling process. Crawling and indexing are separate stages.

Crawled is not indexed

Crawling does not guarantee that a page will be indexed or appear in search results. A page may be successfully crawled but excluded from the index because of duplicate content, indexing directives, or quality-related decisions.

How search engines discover URLs

Search engines use several methods to discover pages. The sources website owners can directly influence are internal links, XML sitemaps, and, to a lesser degree, external links.

Internal links

Internal links connect one page on your website to another page on the same domain. They help users navigate your site and provide crawlers with paths to discover other URLs. For example:

  • Your homepage links to your main service pages.
  • A category page links to individual articles.
  • A blog post links to a related tutorial.
  • A product page links to relevant product categories.

Example: URL discovery through internal links

Homepage/
SEO Fundamentals/academy/seo-fundamentals/
Technical SEO/academy/technical-seo/

Each category can link to its individual lessons.

A page with no internal links pointing to it is often called an orphan page. Such a page may still be discovered through a sitemap or external link, but it has fewer internal discovery paths.

XML sitemaps

An XML sitemap is a file that lists URLs you want search engines to know about. It can be particularly useful for large websites, recently published content, and pages that are not easily reached through internal links. If Search Console reports problems with yours, see these common XML sitemap errors and fixes. A sitemap commonly lives at a URL such as https://www.example.com/sitemap.xml.

A sitemap can help search engines discover URLs, but including a URL does not guarantee that the search engine will crawl or index it.

External links

External links are links from other websites that point to your pages. Search engines can discover your website when their crawlers follow these links. For example, a blog post on another domain might link to one of your tutorials, and a search engine that finds that link may use it as a discovery path to your page.

How crawlers request pages

Once a crawler discovers a URL, it can send an HTTP request to retrieve the page. The web server responds with an HTTP status code and, in many cases, the page's HTML content. Some common HTTP response codes include:

Status codeMeaningCrawling implications
200OKThe server successfully returned the requested page.
301Moved PermanentlyThe URL redirects permanently to another URL.
302FoundThe URL temporarily redirects to another URL.
404Not FoundThe requested page was not found.
403ForbiddenThe server refuses access to the requested resource.
500Internal Server ErrorThe server encountered an error processing the request.
503Service UnavailableThe server is temporarily unable to handle the request.

A successful 200 response does not automatically mean that the page is eligible for indexing. Similarly, a redirect can be a normal part of a website's structure, but unnecessary redirect chains can make crawling less efficient. Mixed protocols are a common cause, which is why it helps to fix HTTPS pages with internal links to HTTP.

Example: What happens when a crawler encounters a redirect?

301 redirectRequested URL/old-page/
200 OKDestination URL/new-page/

A crawler can follow the redirect to retrieve the destination page, subject to its crawling rules and limits.

What is crawl budget?

Crawl budget refers to the amount of crawling a search engine is able and willing to perform on a website over a given period. Google describes it in terms of crawl capacity and crawl demand.

Crawl budget is not usually a concern for small websites with a limited number of useful URLs. It becomes more relevant for very large sites, sites that generate many URLs dynamically, or sites with frequent changes. Factors that can affect crawling include:

  • Server capacity: slow responses and server errors can affect how frequently a crawler can access a site.
  • URL quality: duplicate, parameter-heavy, or low-value URLs can consume crawling resources.
  • Website changes: search engines may revisit URLs more frequently when they expect updates.
  • Crawl demand: search engines may prioritize URLs based on their perceived importance and freshness.

For many websites, improving internal linking, eliminating unnecessary URL variations, and fixing server errors are more practical priorities than trying to optimize an abstract crawl budget number.

What can prevent crawlers from accessing pages?

A crawler may fail to retrieve a page for several technical reasons. Identifying these issues is an important part of technical SEO, and you can check crawlability for Google, Bing, and AI bots to see how they apply to your site.

Robots.txt restrictions

The robots.txt file tells crawlers which URL paths they may access. A disallowed URL may not be crawled by compliant crawlers, although the URL could still be discovered through other sources. To see how AI bots in particular request your pages, you can track AI crawlers with Cloudflare.

Broken internal links

Links pointing to nonexistent or unavailable URLs lead crawlers to error pages instead of useful content. See how to find and fix broken internal links.

Server errors

Repeated 5xx responses can prevent crawlers from retrieving content and may cause them to reduce crawling.

Redirect chains and loops

A chain sends crawlers through multiple redirects before reaching a destination. A redirect loop prevents them from reaching a final page.

Incorrect URL handling

Unnecessary parameters, inconsistent URL formats, and duplicate URLs create extra crawling paths that are not useful to users or search engines. Consistent canonical tags help consolidate those variations.

Crawling vs. indexing

Crawling and indexing are related but distinct processes. Crawling involves discovering and retrieving pages. Indexing involves analyzing and potentially storing their content in a search engine's index.

CrawlingIndexing
Finds and requests URLs.Analyzes and processes retrieved content.
Follows links and uses discovery sources.Evaluates content and canonicalization signals.
Can encounter access restrictions or server errors.Can exclude pages because of directives, duplication, or other factors.
Does not guarantee inclusion in search results.Makes eligible pages available for consideration in search results.

A page may be crawled without being indexed. It is also possible for a URL to be known to a search engine without its content being crawled.

For troubleshooting, first determine whether the URL is accessible to the crawler. Then investigate whether indexing directives, canonical tags, duplication, or other factors affect its eligibility for indexing. The guide to noindex issues covers the most common directive problems.

How to audit crawling with SiteAuditLint

A website crawler such as SiteAuditLint can help you inspect the URLs and links on your site. A crawl provides a structured view of how pages connect, which URLs return errors, and where technical problems may interrupt discovery.

SiteAuditLint crawling workflow

  1. Enter your website URLStart with your homepage or the section of your site you want to audit. Check that the URL uses the correct protocol and hostname.
  2. Configure crawl settingsDecide whether you need to inspect the entire website or a specific section, and make sure your crawl settings match the scope of the audit.
  3. Start the crawlLet the crawler request pages and collect URL and link information. Larger websites may take longer to crawl.
  4. Review crawl resultsExamine the discovered URLs, HTTP status codes, redirects, broken links, and other technical findings.
  5. Fix and compareCorrect issues that affect important pages and links, then run another crawl to compare the new results with the previous audit.

What to look for in a crawl report

  • Discovered URLs: check whether important pages are being found through your site's internal links.
  • Broken links: run a broken link audit to identify internal links that point to missing pages and update them or restore the destination.
  • Redirects: review redirect destinations, chains, and loops.
  • HTTP status codes: find important pages returning errors or unexpected responses.
  • Internal link structure: look for important pages with few internal links or none at all.
  • Crawl comparisons: compare successive audits to see whether issues have been resolved or new ones have appeared.
A crawl report is a snapshot

It shows the website from the crawler's perspective. It is not a record of what Googlebot has crawled or indexed. For Google-specific data, use Google Search Console alongside a technical SEO audit.

Practical exercise: find a page crawlers may struggle to discover

Use this exercise to apply the concepts from the lesson to a real website.

Crawling audit checklist

0 of 6 tasks completed

Key takeaways

  • Crawling is how search engines discover URLs and request page content.
  • Internal links, XML sitemaps, and external links can help crawlers discover pages.
  • HTTP status codes tell crawlers whether a page was retrieved, redirected, or unavailable.
  • Crawl budget is more relevant to large or technically complex websites than to most small sites.
  • Crawling is not the same as indexing. A crawled page is not guaranteed to appear in search results.
  • Regular technical crawls help you find broken links, redirect problems, server errors, and weak internal linking.

Ready to go further? Work through the technical SEO checklist, or compare options in the best desktop SEO crawlers for Windows.