SEO Fundamentals · Lesson 03

Indexing

How search engines process crawled page content, evaluate indexing signals, and decide whether a URL is stored in the search index, plus how to diagnose pages that are missing.

03 Processing10 min Read timeBeginner Level

03 / Processing

Indexing is the process search engines use to analyze and store information about web pages so they can be considered for search results. After a search engine discovers and crawls a URL, it may process the page's content, evaluate its signals, and decide whether to include it in its search index.

This lesson explains what happens during processing and rendering, which indexing signals such as noindex directives, canonical tags, and duplicate content influence index inclusion, and how to use Google Search Console and SiteAuditLint to find out why a page is missing from search results even when visitors can reach it.

What is indexing in SEO?

Indexing is the stage in which a search engine processes information collected from a crawled page and potentially adds it to its searchable database, known as the search index. During processing, search engines analyze different elements of a page, including:

  • Text content: the words, headings, and information presented on the page.
  • HTML structure: elements that help describe the page's content and organization.
  • Links: internal and external links that provide context and connect pages.
  • Metadata: information such as page titles, meta descriptions, and canonical tags.
  • Images and other media: visual content and associated information such as alt text that help search engines understand the page.

A page being crawled does not automatically mean it will be indexed. Search engines determine which pages to include based on their own systems and evaluation processes.

Key concept

Crawling is when a search engine fetches a page. Indexing is when it processes the page and considers whether to store it in its index. Ranking is the process of selecting and ordering results for a search query.

How does indexing work?

Indexing follows crawling in the search engine process, although the stages can overlap and may not happen immediately.

From discovery to search results

01DiscoveryA search engine finds a URL through links, sitemaps, or other sources.
02CrawlingThe crawler requests the URL and retrieves the page.
03ProcessingThe search engine analyzes content, structure, and other signals.
04IndexingThe page may be added to the search index and considered for retrieval.
05RankingThe page may appear in search results when relevant to a query.

Simplified search engine pipeline. Stages can overlap and timing varies.

What happens during processing?

Once a search engine fetches a page, it can analyze the information it receives and determine how that page relates to other pages and topics. For example, a page about technical SEO might contain a title, headings, explanatory text, images, and links to related topics. These elements provide information about the page's subject and structure.

Search engines may also process rendered content, including content generated by JavaScript. The exact processing methods and timing vary by search engine, so important content that only appears after rendering can take longer to be evaluated.

What determines whether a page gets indexed?

There is no single factor that guarantees a page will be indexed. However, several technical and content-related signals can affect whether a search engine can access, process, and select a page for its index.

Crawl accessibility

Search engines need to access a page to crawl and process it. If the page is blocked by robots.txt, requires authentication, or encounters server errors, the search engine may not be able to retrieve its content. Check that important pages are accessible to search engine crawlers and return the expected HTTP status code.

Robots directives

Robots directives tell search engines how to handle a page. A noindex directive tells a search engine not to include a page in its index. It can be implemented using a meta robots tag in the HTML or an X-Robots-Tag HTTP header. For example:

HTML
<meta name="robots" content="noindex">

If a page is blocked from crawling through robots.txt, a search engine may not be able to read its noindex instruction. These are separate mechanisms and should be configured carefully. The guide to noindex issues covers the most common directive problems.

Canonical URLs

Websites can have multiple URLs containing identical or very similar content. Canonical tags help indicate which URL is preferred for indexing. For example, a product page might be accessible through multiple URLs because of tracking parameters or sorting options:

HTML
<link rel="canonical" href="https://example.com/products/shoes/">

Search engines may choose a different canonical URL from the one specified, depending on their assessment of the available signals. Review these common canonical tag problems and fixes if your declared and selected canonicals disagree.

Content quality and uniqueness

Search engines may choose not to index pages that offer little distinct value or contain substantially duplicated content. For example, a website with many pages containing nearly identical descriptions and only minor keyword changes may have difficulty getting all those pages indexed. Create useful, specific content that serves a clear purpose for visitors, and review thin content that adds little on its own.

Internal linking

Internal links help search engines discover pages and understand the relationships between them. Important pages should be accessible through relevant links from other pages on the site. A page with no internal links pointing to it, an orphan page, may be more difficult for crawlers to discover and may receive fewer signals about its importance within the website.

Crawled, currently not indexed

A page can be successfully crawled without being included in a search engine's index. In the Google Search Console page indexing report, the status "Crawled, currently not indexed" means Google crawled the page but did not index it at that time. The status alone does not identify the exact reason. Areas to investigate include:

Duplicate or substantially similar content

Near-identical pages compete with each other, and a search engine may keep only one version.

Limited distinct information

Pages that repeat information found elsewhere offer little reason to be stored separately.

Canonical signals pointing to another URL

A canonical tag, redirect, or internal linking pattern may indicate that a different URL is preferred.

Weak connection to the rest of the site

Pages with few or no internal links receive fewer signals about their importance.

Changes in evaluation

Google may change how it evaluates the page or the site over time.

Review the individual URL and its surrounding technical and content signals before deciding what changes to make.

Crawling vs. indexing vs. ranking

ProcessWhat it meansWhat to check
CrawlingA search engine requests and fetches a page.HTTP status, robots.txt, server access
IndexingA search engine processes the page and considers it for storage in its index.Noindex directives, canonicals, content
RankingA search engine selects and orders results for a particular query.Relevance, content, competition, search context

These are related but separate processes. A page can be crawled but not indexed, or indexed but not appear prominently for a particular search query.

How to check a page's indexing status

Google Search Console provides tools for inspecting individual URLs and reviewing indexing information for a website.

  1. Open URL InspectionSign in to Google Search Console, select the property for your website, and enter the URL you want to inspect into the URL Inspection tool.
  2. Review the reported statusCheck whether Google reports the page as indexed or not indexed. Review any available information about crawling, indexing, and the Google-selected canonical URL.
  3. Investigate technical issuesIf the page is not indexed, check whether the URL returns the expected response, whether robots.txt blocks crawling, whether a noindex directive is present, whether the canonical tag points to the intended URL, whether the page contains useful and distinct content, and whether internal links lead to the page.
  4. Make corrections and monitorIf you identify an issue, correct it and verify that the change is reflected in the live page. You can then use the URL Inspection tool to request indexing when appropriate. A request does not guarantee that Google will crawl or index the page.

How SiteAuditLint helps with indexability checks

A technical SEO audit crawl can help identify website-wide issues that may interfere with search engine access and indexing. With SiteAuditLint, you can inspect technical signals across crawled URLs, including:

  • HTTP status codes: identify pages returning errors or unexpected responses.
  • Robots directives: find pages with noindex instructions in meta robots tags.
  • Canonical tags: review canonical URLs and identify potential configuration issues.
  • Internal links: inspect how pages connect and identify URLs with limited internal linking.
  • Crawl comparisons: compare audit findings across crawls to track changes and confirm whether previously identified issues have been resolved.
Important distinction

SiteAuditLint can identify technical signals and potential indexability issues in a crawl. It cannot confirm whether Google has indexed a particular page. Use Google Search Console for Google's reported indexing status.

Common indexing mistakes

Assuming that a 200 status guarantees indexing

A successful HTTP response only confirms that the server returned a successful response to the request. It does not mean the page has been indexed.

Using robots.txt to remove a page from search results

Robots.txt controls crawling, not guaranteed removal from the index. If you need a page excluded, use the appropriate noindex and removal mechanisms and ensure search engines can access the directive.

Expecting an XML sitemap to guarantee indexing

An XML sitemap helps search engines discover URLs and understand information about them. Submitting a URL does not force a search engine to index it. See these common XML sitemap errors and fixes.

Changing canonical tags without checking the page relationship

Canonical tags should reflect the preferred version of duplicate or substantially similar pages. Incorrect canonical signals can cause search engines to select a different URL than intended.

Requesting indexing without resolving the underlying issue

Repeatedly requesting indexing does not address problems such as accidental noindex directives, duplicate content, or inaccessible pages. Investigate the reported status and correct any relevant issues first.

Knowledge check

Test what you learned about crawling, indexing directives, and canonicalization.

1. Does crawling guarantee that a page will be indexed?
2. What does a noindex directive tell a search engine?
3. What is the purpose of a canonical tag?

Key takeaways

  • Crawling and indexing are separate stages in the search engine process.
  • Search engines process page content and technical signals before deciding whether to include a page in their index.
  • Robots directives, canonical tags, content uniqueness, accessibility, and internal links are useful areas to investigate when diagnosing indexing issues.
  • A successful crawl, a sitemap submission, or an indexing request does not guarantee inclusion in search results.
  • Use a technical crawler to inspect site-wide signals and Google Search Console to investigate Google's reported indexing status.

Ready to go further? Work through the technical SEO checklist to review indexability alongside the rest of your site.