Home›Academy›Technical SEO›XML sitemaps

Technical SEO · Lesson 05

XML sitemaps

How an XML sitemap lists the canonical, indexable URLs you want discovered, and how sitemap errors can send mixed signals to search engine crawlers.

05 Technical SEO13 min Read timeBeginner Level

05 / Technical SEO

An XML sitemap is a file that gives search engines a list of the URLs you consider important to discover and crawl.

A sitemap does not force Google to index a URL. It is a discovery signal, a hint about which canonical pages matter on your website. When the URLs inside the sitemap return 200 OK, are indexable, and match their canonical tags, Googlebot gets a clean, trustworthy list. When the sitemap is full of redirects, 404s, and noindex pages, the signal becomes noisy and less useful.

This lesson covers sitemap URLs, sitemap index files, lastmod dates, common sitemap errors, and non-indexable URLs. It also explains how to check sitemap consistency during a technical SEO audit and use SiteAuditLint to find sitemap URLs that do not behave like clean indexable pages.

1. What is an XML sitemap?

An XML sitemap follows the sitemaps.org protocol. A basic sitemap looks like this:

<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">

  <url>
    <loc>https://example.com/</loc>
  </url>

  <url>
    <loc>https://example.com/services/</loc>
  </url>

  <url>
    <loc>https://example.com/blog/</loc>
  </url>

</urlset>

Each <url> entry contains a URL you want search engines to discover. The <urlset> element wraps the whole list and declares the sitemap namespace.

What belongs in a sitemap?

A sitemap should generally contain URLs that are:

  • Canonical, meaning the preferred version of the page.
  • Indexable, with no noindex directive.
  • Successful, returning a 200 response.
  • Important to the site and its visitors.
  • Accessible to search engines and not blocked by robots.txt.
A sitemap is not a URL dump

Avoid treating the sitemap as a complete list of every URL a crawler can technically find. Parameter URLs, filtered views, internal search results, and staging paths do not belong there. The goal is a clean list of URLs you want discovered and considered for indexing.

How a sitemap supports discovery

SubmitSitemap foundVia Google Search Console, robots.txt, or a known path like /sitemap.xml.
ReadURLs parsedThe crawler reads each <loc> and any <lastmod> value.
CrawlPages consideredListed URLs are crawled and evaluated alongside canonicals, links, and content.

Discovery through the sitemap is a hint. Indexing still depends on the page itself.

2. Sitemap URLs

The <loc> element contains the absolute URL the sitemap recommends to search engines:

<url>
  <loc>https://example.com/products/widget/</loc>
</url>

The URL should match the canonical version of the page exactly, including protocol, subdomain, and trailing slash. If https://example.com/page redirects to https://example.com/page/, the sitemap should contain the final canonical URL rather than the redirecting one.

Check your sitemap URLs

Every listed URL should:

  • Return 200 OK.
  • Be indexable.
  • Use HTTPS when HTTPS is the canonical version.
  • Match the preferred canonical URL.
  • Not redirect.
  • Not be blocked from crawling.
  • Not be a duplicate version of another URL.

Example

Bad sitemapBetter sitemap
https://example.com/old-pagehttps://example.com/
https://example.com/page?sort=pricehttps://example.com/products/
https://example.com/private/https://example.com/products/widget/
https://example.com/missing-page

The first list mixes a redirected URL, a sorting parameter, a private path, and a missing page. The second gives crawlers a clearer set of canonical URLs to discover.

3. Sitemap index files

A single sitemap file can hold up to 50,000 URLs or 50 MB uncompressed. Large websites split their URLs across several files and reference them from a sitemap index file:

<?xml version="1.0" encoding="UTF-8"?>
<sitemapindex xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">

  <sitemap>
    <loc>https://example.com/sitemap-pages.xml</loc>
  </sitemap>

  <sitemap>
    <loc>https://example.com/sitemap-products.xml</loc>
  </sitemap>

  <sitemap>
    <loc>https://example.com/sitemap-blog.xml</loc>
  </sitemap>

</sitemapindex>

The index points to individual sitemap files, and a site might organize them by content type:

Sitemap index structure

indexSitemap index/sitemap.xml
xmlStatic pages/sitemap-pages.xml
xmlProduct pages/sitemap-products.xml
xmlBlog posts/sitemap-blog.xml
xmlCategory pages/sitemap-categories.xml

Splitting by content type also makes Google Search Console reporting more useful, because you can see indexing coverage for products separately from blog posts.

What to check

Index accessibility

The sitemap index itself returns 200 and parses as valid XML.

Referenced files exist

Every child sitemap listed in the index exists and returns a successful response.

Valid child URLs

URLs inside each child sitemap are canonical, indexable, and live.

Retired files removed

Old sitemap files are removed from the index when they are no longer used.

A sitemap index containing broken sitemap references creates unnecessary discovery problems.

4. lastmod dates

The optional <lastmod> element tells search engines when a URL was last significantly modified, using W3C Datetime format:

<url>
  <loc>https://example.com/blog/seo-basics/</loc>
  <lastmod>2026-09-30</lastmod>
</url>

The date should represent a meaningful change to the page, such as updated main content, structured data, or links, rather than the date the sitemap file was generated. Google uses lastmod when it is consistently and verifiably accurate, and ignores it when it is not. The <changefreq> and <priority> elements are ignored by Google, so accurate lastmod values carry far more weight than either.

Noisy dates lose trust

Stamping today's date on every URL each time the sitemap regenerates, without changing the content, trains search engines to disregard your lastmod values entirely.

Check your dates

  • Missing dates where meaningful update dates are available.
  • Invalid date formats.
  • Future dates.
  • Dates that change without corresponding content changes.
  • The same update date applied automatically to every URL.

Accurate lastmod values help search engines prioritize recrawling the URLs that have actually changed.

5. Sitemap errors

Sitemap errors make the file less useful as a discovery source. In Google Search Console, they often surface in the Sitemaps report as statuses like "Couldn't fetch" or "Sitemap could not be read".

404 sitemap

https://example.com/sitemap.xml returns 404 Not Found, so search engines cannot retrieve the sitemap at all.

Invalid XML

Malformed XML, unescaped characters such as a raw &, or incomplete tags can prevent the sitemap from being processed.

Broken sitemap references

An index lists https://example.com/sitemap-products.xml when that file no longer exists.

Incorrect URLs

The sitemap is reachable, but the URLs inside it redirect, return 404 or 5xx, are blocked, carry noindex, or canonicalize elsewhere.

Here is an invalid XML example where the closing tag is incomplete:

<url>
  <loc>https://example.com/page</loc>
</url

The sitemap file itself may be accessible while the URLs inside it are problematic, so a successful fetch is only the first check.

6. Non-indexable URLs

One of the most important sitemap checks is whether the listed URLs are actually eligible for indexing. Consider a sitemap that lists https://example.com/product-a/ while the page contains:

<meta name="robots" content="noindex">

The sitemap is recommending a URL that the page itself says should not be indexed. Similarly, if the sitemap lists https://example.com/old-page/ and that URL returns 301 to https://example.com/new-page/, the sitemap should be updated to reference the canonical destination.

Common non-indexable sitemap URLs

ProblemWhat it means
noindexPage explicitly excluded from indexing
RedirectURL redirects to another URL
404Page no longer exists
410Page has been permanently removed
Canonical mismatchPage points to another canonical URL
Blocked URLRobots.txt prevents crawlers from accessing the URL
Server errorURL returns a 5xx response

For more on how these responses behave, see Lesson 04: Status codes and Lesson 02: Canonicals.

7. Sitemap consistency

A sitemap works best when its signals agree with the rest of the website. Compare two setups for https://example.com/services/seo/:

SignalConsistent setupConflicting setup
Sitemap/services/seo//services/seo/
Canonical/services/seo//services/
HTTP response200 OK200 OK
Robots metaindexnoindex

In the conflicting setup, the sitemap says one thing while the canonical tag and the indexing directive say something else. Search engines have to resolve the contradiction themselves, and the sitemap's recommendation loses credibility. When the sitemap, canonical, status code, robots meta, and internal links all point to the same URL, the signal is at its strongest.

8. How to audit sitemaps with SiteAuditLint

A website crawler such as SiteAuditLint helps you identify URLs that appear in your sitemap but do not behave like clean indexable pages.

SiteAuditLint sitemap workflow

  1. Find the sitemapLocate /sitemap.xml or the sitemap referenced by the Sitemap: line in your site's robots.txt.
  2. Crawl the siteRun a crawl with SiteAuditLint so sitemap URLs can be compared with what the crawler finds.
  3. Check sitemap URLsReview URLs discovered through the sitemap and compare their HTTP responses, indexability, canonical URLs, and redirects.
  4. Find mismatchesFlag sitemap URLs that return errors, redirect, are non-indexable, have canonical mismatches, or no longer exist.
  5. Clean the sitemapRemove URLs that should not be included and replace outdated URLs with their current canonical versions.
  6. RecheckRun another crawl after the sitemap has been updated, then resubmit it in Google Search Console.
Reliable beats complete

The goal is not to put every known URL into the sitemap. It is to maintain a reliable list of URLs you want search engines to discover. Also compare in the other direction: important indexable pages missing from the sitemap are worth adding.

9. Practical example

Imagine a website has 1,000 URLs in its sitemap. After crawling them, you find:

Sitemap URLsCountAction
Valid, indexable URLs870Keep
Redirects45Replace with final URLs
404 errors30Remove if permanently gone
noindex URLs25Remove if intentionally excluded
Canonical mismatches20Use the canonical URL instead
5xx errors10Investigate the underlying problem
Total1,000

The sitemap technically contains 1,000 URLs, but only 870 are clean indexable targets. The remaining 130 need investigation. After cleanup, the sitemap provides a much clearer discovery signal, and the gap between "submitted" and "indexed" in the Search Console Sitemaps report becomes easier to interpret.

10. Sitemap checklist

Before considering an XML sitemap clean, work through this list.

XML sitemap audit checklist

0 of 15 tasks completed

11. Key takeaways

  • An XML sitemap is a discovery signal, not an indexing command.
  • A useful sitemap contains the canonical, indexable URLs you want search engines to discover.
  • Sitemap URLs should match the canonical version exactly and return 200 OK.
  • Sitemap index files should reference only existing, current sitemap files.
  • lastmod dates should reflect meaningful content changes, and Google ignores changefreq and priority.
  • Sitemap errors such as a 404 sitemap or invalid XML prevent reliable processing.
  • Non-indexable URLs such as redirects, 404s, noindex pages, and canonical mismatches should not be recommended for discovery.
  • A clean sitemap keeps discovery signals consistent with canonicals, status codes, and robots directives.

Knowledge check

1. What does including a URL in an XML sitemap do?
2. A sitemap lists /old-page/, which returns 301 to /new-page/. What should you do?
3. Your CMS updates every URL's lastmod to today whenever the sitemap regenerates. What is the problem?
4. A sitemap URL returns 200 OK but its canonical points to another page. What is this?

Ready to go further? Review the technical SEO checklist, or revisit Lesson 04: Status codes.