Crawlability vs Indexability: What’s the Difference?

AI OVERVIEW

Crawlability determines whether search engines can access a page; indexability determines whether they can include it in search results.

You run a technical SEO audit and tell a client their pages are being crawled. The response comes quickly: "So they're indexed, right? And if they're indexed, they should rank."

Not quite.

Crawling, indexing, and ranking are three different stages of how search engines handle a page. Googlebot can crawl a URL without indexing it. A page can be indexed without ranking well. And a page that isn't being crawled properly has a problem much earlier in the process.

The confusion is common, especially with clients who already consider themselves SEO experts. To someone who treats crawling and indexing as the same thing, "the page was crawled" sounds like the job is done. It isn't.

If you regularly need to explain why a crawled page isn't necessarily an indexed page, or why being indexed doesn't guarantee rankings, the distinction needs to be clear. Here's how to separate the two, diagnose the difference, and explain it in a way even the "I already know SEO" client can follow.

Crawlability vs Indexability at a Glance

Each stage answers a different question, and a "yes" to one says nothing about the next.

Gate 1

Crawlability

Can Google get to the page?

Discovery and access. Controlled by robots.txt, internal links, and HTTP responses.

Gate 2

Indexability

Can Google use the page in its index?

Eligibility. Controlled by noindex, canonical tags, and HTTP responses.

Gate 3

Ranking

How does the page perform in search?

Visibility. Depends on relevance, quality, and competition.

CrawlabilityIndexability
Main questionCan a search engine access the URL?Can the URL be included in the search index?
Main concernDiscovery and accessEligibility for indexing
Common controlsrobots.txt, links, HTTP responsesnoindex, canonical, HTTP responses
Typical problemURL cannot be crawledURL can be crawled but should not or may not be indexed
Audit focusCan the crawler reach the page?What happens after the page is accessed?

What Crawlability Means

Crawlability describes whether search engine crawlers (see how a website crawler works) can discover a URL and access its content. To crawl a page, a crawler generally works through five steps:

  1. Discover the URL.
  2. Request the URL.
  3. Receive an HTTP response.
  4. Access the content and resources needed to process and render the page.
  5. Follow additional links when appropriate.

A URL can become difficult or impossible to crawl when:

  • robots.txt blocks the URL
  • The server returns an error
  • It redirects through a problematic chain
  • No useful internal links point to it
  • Discovery depends on hard-to-process mechanisms
  • The server times out
  • Important resources are inaccessible
  • The URL requires authentication
In shortCrawlability is about access and discovery.

What Indexability Means

Indexability describes whether a URL can be considered for inclusion in a search engine's index. A page can be perfectly accessible and still tell search engines not to index it:

<meta name="robots" content="noindex">

The crawler reaches the page, reads the HTML, and sees the noindex directive. Here's what that URL looks like in an audit:

CheckResult
URL discoverableYes
URL crawlableYes
HTTP response200
Blocked by robots.txtNo
noindex presentYes
IndexableNo

The page is crawlable but not indexable, one of the most common noindex issues. That's the distinction clients most often miss.

Crawled Does Not Mean Indexed

This causes so much confusion in client conversations that it's worth stating plainly. A crawler visiting a page only means the search engine was able to access it. It does not mean:

  • The page was added to the index
  • The page will appear in search results
  • The page will rank
  • The page will rank well
  • It's the preferred version of a duplicate URL

Think of it as a pipeline where a URL can stall at any stage:

Crawling is a step in the process, not the finish line.

This is also why Google Search Console has a status called "Crawled, currently not indexed." Google reached the page and decided not to index it, at least for now.

Crawlable but Not Indexable

Consider https://example.com/products/widget. The server returns 200 OK, the page is linked internally, and robots.txt doesn't block it. The crawler reaches it without trouble. But the HTML contains <meta name="robots" content="noindex">.

Crawlable: Yes

Linked internally, returns 200, not blocked.

Indexable: No

The page carries an instruction not to index it.

A crawl report showing a URL was successfully crawled should never be read as proof that the URL is indexed.

Indexable but Difficult to Crawl

The opposite happens too. A URL might pass every indexability check:

  • Returns 200 OK
  • No noindex
  • Self-referencing canonical
  • Not blocked by robots.txt

But suppose the only way to reach it is through a complicated JavaScript interaction, with no crawlable internal links pointing to it. The URL may be eligible for indexing while its discovery and crawl path are weak.

The takeawayCrawlability and indexability need to be audited separately. Passing one tells you nothing about the other.

robots.txt

robots.txt controls crawler access to URL paths:

User-agent: *
Disallow: /private/

A URL under /private/ may be blocked from crawling. But robots.txt is a crawl control, not an indexing directive. A blocked URL can still be discovered through links or other references and may appear in search results without a description. That's why robots.txt and noindex are not interchangeable.

Common trapIf you block a URL in robots.txt, Google can't crawl it, so it never sees a noindex on that page. Use one or the other deliberately, not both on the same URL.

Audit checks for robots.txt:

  1. Is the URL blocked?
  2. Which User-agent rule applies?
  3. Is an entire directory blocked?
  4. Is an important page accidentally covered by a broad rule?
  5. Are important CSS, JavaScript, or image resources unnecessarily blocked?

Meta Robots

The meta robots element sits in the HTML <head> and tells crawlers how the page should be handled. For crawlability vs indexability, noindex is the directive that matters most.

1

Crawlability question

Can the crawler reach the page?

2

Indexability question

What instructions does the crawler receive after reaching it?

X-Robots-Tag

The X-Robots-Tag delivers robots directives through the HTTP response header instead of the HTML:

HTTP/1.1 200 OK
Content-Type: text/html
X-Robots-Tag: noindex

The page can look completely normal in a browser while the response header tells crawlers not to index it. It's especially worth checking on:

  • PDFs
  • Images
  • Downloadable files
  • Non-HTML resources
  • Server-level SEO rules
Audit tipA complete indexability audit checks both the HTML robots directives and the HTTP headers.

Canonicals

Canonical tags help search engines understand which URL should represent a group of duplicate or substantially similar URLs:

<link rel="canonical" href="https://example.com/widget">

A canonical is not the same as noindex. A URL can be fully crawlable while naming another URL as its canonical.

Check whether the canonical (and watch for these canonical tag issues):

  • Exists
  • Uses a valid URL
  • Returns a successful response
  • Points to the intended URL
  • Avoids unnecessary redirects
  • Matches internal linking
  • Isn't pointing to an unrelated page

HTTP Responses

Status codes shape both sides of the picture.

ResponseWhat it generally indicates
200Page successfully returned
301Permanent redirect
302Temporary redirect
403Access forbidden
404Resource not found
410Resource permanently gone
5xxServer-side error

A 200 OK does not mean a page will be indexed, and a URL returning an error can't be treated as normally accessible content. When redirects are involved, evaluate the final destination separately. Follow this path for every URL:

Internal Links

Internal links are one of the strongest crawl discovery signals. When important pages are linked through categories, navigation, breadcrumbs, and contextual links, crawlers have multiple paths to find them. Pages with no internal links pointing to them are orphan pages, and they're much harder to discover through the site's normal architecture.

  • Homepage
    • Category
      • Subcategory
        • Product
          • Related product

Each internal link creates another discovery path.

Audit questions:

  1. Is there at least one internal link to the URL?
  2. Can it be reached from important pages?
  3. Are links using valid URLs?
  4. Are links pointing to redirects?
  5. Are important pages buried several clicks deep?
  6. Are links implemented in a way crawlers can process, such as standard <a href> elements?

Crawl Discovery

Crawling starts with discovery. Search engines find URLs through several sources:

  • Internal links
  • XML sitemaps
  • External links (backlinks)
  • Previously known URLs
  • Redirects
  • Other page references

An XML sitemap is useful, but it does not replace good internal linking.

Weaksitemap.xml → /products/widget

Listed in the sitemap, but no important pages link to it. Little context or discovery support.

StrongHomepage → Products → Widget

Linked through the architecture and also listed in sitemap.xml. Sources reinforce each other.

How to Diagnose the Difference

When a client asks, "The page was crawled, so why isn't it ranking?", don't jump straight to ranking. Start with the URL and work forward through each gate.

Step by step

  1. 1Check discoveryCan the URL be found through internal links, the XML sitemap, or other known sources?
  2. 2Check robots.txtIf blocked, check access rules. If not, continue.
  3. 3Check the HTTP responseLook for 200, 301, 302, 403, 404, 410, or 5xx. Follow any redirect chain to the final destination.
  4. 4Check robots directivesInspect both <meta name="robots"> and the X-Robots-Tag header.
  5. 5Check the canonicalA canonical pointing elsewhere isn't automatically broken, but it changes how the URL should be interpreted.
  6. 6Check internal linksDo important pages actually link to the URL?
  7. 7Classify the problemAssign each finding to the right area so the fix is precise.

Classify every finding

FindingPrimary area
URL blocked by robots.txtCrawlability
No internal discovery pathCrawl discovery
403 responseCrawlability
5xx responseCrawlability
noindexIndexability
X-Robots-Tag: noindexIndexability
Canonical points elsewhereCanonical / indexing signal
Redirect chainCrawl and URL handling
404 / 410URL availability
Valid 200 page with weak discoveryCrawl discovery

Practical Crawlability and Indexability Audit Workflow

Use this workflow when auditing a full site. If AI visibility matters too, also check crawlability for Google, Bing, and AI bots.

1. Start with the URL list

Collect URLs from every source you have, and don't assume the sitemap contains every URL that matters:

  • XML sitemaps
  • Internal links
  • Existing crawl data
  • Analytics and site inventories
  • Known landing pages

2. Test URL accessibility

Record the URL, HTTP status, final URL, redirect count, and response time. Then flag:

  • 4xx responses
  • 5xx responses
  • Excessive redirects
  • Redirect loops
  • Unexpected destinations

3. Check robots.txt

Record each URL as Allowed or Blocked, paying particular attention to important pages accidentally caught by broad rules.

4. Check indexability directives

Inspect both <meta name="robots"> and the X-Robots-Tag header, and flag any unexpected noindex.

5. Check canonicalization

For each URL, record:

  • Canonical URL
  • Self-canonical?
  • Canonical status code
  • Canonical redirects?
  • Canonical blocked?

Then compare the canonical with the URL's intended role.

6. Check internal links

For important URLs, map which internal links point to them and flag:

  • Orphan pages
  • Broken internal links
  • Links to redirects
  • Links to 4xx URLs
  • Inconsistent canonical targets
  • Excessive click depth

7. Separate the findings

Don't lump everything into a single "indexing issue" bucket. Three clear categories make the audit report far easier to understand and act on:

Crawlability

  • robots.txt
  • HTTP errors
  • Redirects
  • Server access

Discovery

  • Internal links
  • XML sitemap
  • Orphan URLs
  • Crawl depth

Indexability

  • Meta robots
  • X-Robots-Tag
  • Canonical
  • Indexing signals

A Simple Diagnostic Matrix

CrawlableIndexability statusWhat it means
YesEligibleURL can be accessed and has no obvious blocking index directive
YesNot eligibleCrawler can access it, but an indexing control or signal prevents normal indexing
NoCannot fully evaluateThe crawler can't properly access the URL, so full page-level indexability analysis isn't possible
DifficultTechnically eligibleURL may be eligible for indexing, but discovery or access needs attention
Order mattersCrawlability comes before meaningful indexability analysis. If a crawler can't access a page, you can't reliably evaluate what's happening inside it.

Where SiteAuditLint Fits

A technical SEO crawler brings these checks together so you aren't inspecting URLs one by one. A useful audit moves through every layer in sequence:

The goal isn't simply a list of URLs. It's an explanation of why each URL has its particular crawl or index status, which matters most when a crawl contains hundreds or thousands of URLs.

Find the Problem Before You Explain the Fix

Diagnosing this manually means checking status codes, redirects, robots.txt rules, robots directives, canonicals, internal links, and more for every URL. SiteAuditLint brings those checks together in one technical SEO audit. It points you to the issues affecting your URLs, shows where crawlers are blocked or redirected, identifies indexability signals, flags canonical problems, and surfaces internal linking and crawl discovery issues.

Instead of telling a client "the page was crawled, so it should be indexed," you can show them what actually happened to the URL and where the problem is.

Key Takeaways

Crawlability

Can search engines discover and access the URL?

Indexability

Can the URL be considered for inclusion in the index?

Crawled ≠ indexed

Access is not inclusion.

Indexed ≠ ranked

Inclusion is not visibility.

  • A page can be crawlable but not indexable
  • A page can be eligible for indexing but poorly discovered
  • robots.txt primarily controls crawling
  • meta robots and X-Robots-Tag carry indexing directives like noindex
  • Canonicals signal the preferred URL among related URLs
  • HTTP responses decide whether and how a URL can be accessed
  • Internal links are key paths for URL discovery
  • Audit crawlability, discovery, indexability, and ranking separately

The easiest mental model:

A problem at any stage produces a different SEO outcome. So when a client says, "It was crawled, so why isn't it ranking?", you now have a much better answer:

Because being crawled is only one step in the journey from URL discovery to search visibility.