You run a technical SEO audit and tell a client their pages are being crawled. The response comes quickly: "So they're indexed, right? And if they're indexed, they should rank."
Not quite.
Crawling, indexing, and ranking are three different stages of how search engines handle a page. Googlebot can crawl a URL without indexing it. A page can be indexed without ranking well. And a page that isn't being crawled properly has a problem much earlier in the process.
The confusion is common, especially with clients who already consider themselves SEO experts. To someone who treats crawling and indexing as the same thing, "the page was crawled" sounds like the job is done. It isn't.
If you regularly need to explain why a crawled page isn't necessarily an indexed page, or why being indexed doesn't guarantee rankings, the distinction needs to be clear. Here's how to separate the two, diagnose the difference, and explain it in a way even the "I already know SEO" client can follow.
Crawlability vs Indexability at a Glance
Each stage answers a different question, and a "yes" to one says nothing about the next.
Crawlability
Can Google get to the page?
Discovery and access. Controlled by robots.txt, internal links, and HTTP responses.
Indexability
Can Google use the page in its index?
Eligibility. Controlled by noindex, canonical tags, and HTTP responses.
Ranking
How does the page perform in search?
Visibility. Depends on relevance, quality, and competition.
| Crawlability | Indexability | |
|---|---|---|
| Main question | Can a search engine access the URL? | Can the URL be included in the search index? |
| Main concern | Discovery and access | Eligibility for indexing |
| Common controls | robots.txt, links, HTTP responses | noindex, canonical, HTTP responses |
| Typical problem | URL cannot be crawled | URL can be crawled but should not or may not be indexed |
| Audit focus | Can the crawler reach the page? | What happens after the page is accessed? |
What Crawlability Means
Crawlability describes whether search engine crawlers (see how a website crawler works) can discover a URL and access its content. To crawl a page, a crawler generally works through five steps:
- Discover the URL.
- Request the URL.
- Receive an HTTP response.
- Access the content and resources needed to process and render the page.
- Follow additional links when appropriate.
A URL can become difficult or impossible to crawl when:
robots.txtblocks the URL- The server returns an error
- It redirects through a problematic chain
- No useful internal links point to it
- Discovery depends on hard-to-process mechanisms
- The server times out
- Important resources are inaccessible
- The URL requires authentication
What Indexability Means
Indexability describes whether a URL can be considered for inclusion in a search engine's index. A page can be perfectly accessible and still tell search engines not to index it:
<meta name="robots" content="noindex">
The crawler reaches the page, reads the HTML, and sees the noindex directive. Here's what that URL looks like in an audit:
| Check | Result |
|---|---|
| URL discoverable | Yes |
| URL crawlable | Yes |
| HTTP response | 200 |
| Blocked by robots.txt | No |
noindex present | Yes |
| Indexable | No |
The page is crawlable but not indexable, one of the most common noindex issues. That's the distinction clients most often miss.
Crawled Does Not Mean Indexed
This causes so much confusion in client conversations that it's worth stating plainly. A crawler visiting a page only means the search engine was able to access it. It does not mean:
- The page was added to the index
- The page will appear in search results
- The page will rank
- The page will rank well
- It's the preferred version of a duplicate URL
Think of it as a pipeline where a URL can stall at any stage:
Crawling is a step in the process, not the finish line.
This is also why Google Search Console has a status called "Crawled, currently not indexed." Google reached the page and decided not to index it, at least for now.
Crawlable but Not Indexable
Consider https://example.com/products/widget. The server returns 200 OK, the page is linked internally, and robots.txt doesn't block it. The crawler reaches it without trouble. But the HTML contains <meta name="robots" content="noindex">.
Linked internally, returns 200, not blocked.
The page carries an instruction not to index it.
A crawl report showing a URL was successfully crawled should never be read as proof that the URL is indexed.
Indexable but Difficult to Crawl
The opposite happens too. A URL might pass every indexability check:
- Returns
200 OK - No
noindex - Self-referencing canonical
- Not blocked by
robots.txt
But suppose the only way to reach it is through a complicated JavaScript interaction, with no crawlable internal links pointing to it. The URL may be eligible for indexing while its discovery and crawl path are weak.
robots.txt
robots.txt controls crawler access to URL paths:
User-agent: * Disallow: /private/
A URL under /private/ may be blocked from crawling. But robots.txt is a crawl control, not an indexing directive. A blocked URL can still be discovered through links or other references and may appear in search results without a description. That's why robots.txt and noindex are not interchangeable.
noindex on that page. Use one or the other deliberately, not both on the same URL.Audit checks for robots.txt:
- Is the URL blocked?
- Which
User-agentrule applies? - Is an entire directory blocked?
- Is an important page accidentally covered by a broad rule?
- Are important CSS, JavaScript, or image resources unnecessarily blocked?
Meta Robots
The meta robots element sits in the HTML <head> and tells crawlers how the page should be handled. For crawlability vs indexability, noindex is the directive that matters most.
Crawlability question
Can the crawler reach the page?
Indexability question
What instructions does the crawler receive after reaching it?
X-Robots-Tag
The X-Robots-Tag delivers robots directives through the HTTP response header instead of the HTML:
HTTP/1.1 200 OK Content-Type: text/html X-Robots-Tag: noindex
The page can look completely normal in a browser while the response header tells crawlers not to index it. It's especially worth checking on:
- PDFs
- Images
- Downloadable files
- Non-HTML resources
- Server-level SEO rules
Canonicals
Canonical tags help search engines understand which URL should represent a group of duplicate or substantially similar URLs:
<link rel="canonical" href="https://example.com/widget">
A canonical is not the same as noindex. A URL can be fully crawlable while naming another URL as its canonical.
https://example.com/widget?color=redStill crawlable. Signals another URL as the preferred version.
https://example.com/widgetCheck whether the canonical (and watch for these canonical tag issues):
- Exists
- Uses a valid URL
- Returns a successful response
- Points to the intended URL
- Avoids unnecessary redirects
- Matches internal linking
- Isn't pointing to an unrelated page
HTTP Responses
Status codes shape both sides of the picture.
| Response | What it generally indicates |
|---|---|
| 200 | Page successfully returned |
| 301 | Permanent redirect |
| 302 | Temporary redirect |
| 403 | Access forbidden |
| 404 | Resource not found |
| 410 | Resource permanently gone |
| 5xx | Server-side error |
A 200 OK does not mean a page will be indexed, and a URL returning an error can't be treated as normally accessible content. When redirects are involved, evaluate the final destination separately. Follow this path for every URL:
Internal Links
Internal links are one of the strongest crawl discovery signals. When important pages are linked through categories, navigation, breadcrumbs, and contextual links, crawlers have multiple paths to find them. Pages with no internal links pointing to them are orphan pages, and they're much harder to discover through the site's normal architecture.
- Homepage
- Category
- Subcategory
- Product
- Related product
- Product
- Subcategory
- Category
Each internal link creates another discovery path.
Audit questions:
- Is there at least one internal link to the URL?
- Can it be reached from important pages?
- Are links using valid URLs?
- Are links pointing to redirects?
- Are important pages buried several clicks deep?
- Are links implemented in a way crawlers can process, such as standard
<a href>elements?
Crawl Discovery
Crawling starts with discovery. Search engines find URLs through several sources:
- Internal links
- XML sitemaps
- External links (backlinks)
- Previously known URLs
- Redirects
- Other page references
An XML sitemap is useful, but it does not replace good internal linking.
sitemap.xml → /products/widgetListed in the sitemap, but no important pages link to it. Little context or discovery support.
Homepage → Products → WidgetLinked through the architecture and also listed in sitemap.xml. Sources reinforce each other.
How to Diagnose the Difference
When a client asks, "The page was crawled, so why isn't it ranking?", don't jump straight to ranking. Start with the URL and work forward through each gate.
Step by step
- 1Check discoveryCan the URL be found through internal links, the XML sitemap, or other known sources?
- 2Check robots.txtIf blocked, check access rules. If not, continue.
- 3Check the HTTP responseLook for 200, 301, 302, 403, 404, 410, or 5xx. Follow any redirect chain to the final destination.
- 4Check robots directivesInspect both
<meta name="robots">and theX-Robots-Tagheader. - 5Check the canonicalA canonical pointing elsewhere isn't automatically broken, but it changes how the URL should be interpreted.
- 6Check internal linksDo important pages actually link to the URL?
- 7Classify the problemAssign each finding to the right area so the fix is precise.
Classify every finding
| Finding | Primary area |
|---|---|
URL blocked by robots.txt | Crawlability |
| No internal discovery path | Crawl discovery |
403 response | Crawlability |
5xx response | Crawlability |
noindex | Indexability |
X-Robots-Tag: noindex | Indexability |
| Canonical points elsewhere | Canonical / indexing signal |
| Redirect chain | Crawl and URL handling |
404 / 410 | URL availability |
Valid 200 page with weak discovery | Crawl discovery |
Practical Crawlability and Indexability Audit Workflow
Use this workflow when auditing a full site. If AI visibility matters too, also check crawlability for Google, Bing, and AI bots.
1. Start with the URL list
Collect URLs from every source you have, and don't assume the sitemap contains every URL that matters:
- XML sitemaps
- Internal links
- Existing crawl data
- Analytics and site inventories
- Known landing pages
2. Test URL accessibility
Record the URL, HTTP status, final URL, redirect count, and response time. Then flag:
- 4xx responses
- 5xx responses
- Excessive redirects
- Redirect loops
- Unexpected destinations
3. Check robots.txt
Record each URL as Allowed or Blocked, paying particular attention to important pages accidentally caught by broad rules.
4. Check indexability directives
Inspect both <meta name="robots"> and the X-Robots-Tag header, and flag any unexpected noindex.
5. Check canonicalization
For each URL, record:
- Canonical URL
- Self-canonical?
- Canonical status code
- Canonical redirects?
- Canonical blocked?
Then compare the canonical with the URL's intended role.
6. Check internal links
For important URLs, map which internal links point to them and flag:
- Orphan pages
- Broken internal links
- Links to redirects
- Links to 4xx URLs
- Inconsistent canonical targets
- Excessive click depth
7. Separate the findings
Don't lump everything into a single "indexing issue" bucket. Three clear categories make the audit report far easier to understand and act on:
Crawlability
- robots.txt
- HTTP errors
- Redirects
- Server access
Discovery
- Internal links
- XML sitemap
- Orphan URLs
- Crawl depth
Indexability
- Meta robots
- X-Robots-Tag
- Canonical
- Indexing signals
A Simple Diagnostic Matrix
| Crawlable | Indexability status | What it means |
|---|---|---|
| Yes | Eligible | URL can be accessed and has no obvious blocking index directive |
| Yes | Not eligible | Crawler can access it, but an indexing control or signal prevents normal indexing |
| No | Cannot fully evaluate | The crawler can't properly access the URL, so full page-level indexability analysis isn't possible |
| Difficult | Technically eligible | URL may be eligible for indexing, but discovery or access needs attention |
Where SiteAuditLint Fits
A technical SEO crawler brings these checks together so you aren't inspecting URLs one by one. A useful audit moves through every layer in sequence:
The goal isn't simply a list of URLs. It's an explanation of why each URL has its particular crawl or index status, which matters most when a crawl contains hundreds or thousands of URLs.
Find the Problem Before You Explain the Fix
Diagnosing this manually means checking status codes, redirects, robots.txt rules, robots directives, canonicals, internal links, and more for every URL. SiteAuditLint brings those checks together in one technical SEO audit. It points you to the issues affecting your URLs, shows where crawlers are blocked or redirected, identifies indexability signals, flags canonical problems, and surfaces internal linking and crawl discovery issues.
Instead of telling a client "the page was crawled, so it should be indexed," you can show them what actually happened to the URL and where the problem is.
Key Takeaways
Crawlability
Can search engines discover and access the URL?
Indexability
Can the URL be considered for inclusion in the index?
Crawled ≠ indexed
Access is not inclusion.
Indexed ≠ ranked
Inclusion is not visibility.
- A page can be crawlable but not indexable
- A page can be eligible for indexing but poorly discovered
robots.txtprimarily controls crawlingmeta robotsandX-Robots-Tagcarry indexing directives likenoindex- Canonicals signal the preferred URL among related URLs
- HTTP responses decide whether and how a URL can be accessed
- Internal links are key paths for URL discovery
- Audit crawlability, discovery, indexability, and ranking separately
The easiest mental model:
A problem at any stage produces a different SEO outcome. So when a client says, "It was crawled, so why isn't it ranking?", you now have a much better answer:
Because being crawled is only one step in the journey from URL discovery to search visibility.