05 / Technical SEO
An XML sitemap is a file that gives search engines a list of the URLs you consider important to discover and crawl.
A sitemap does not force Google to index a URL. It is a discovery signal, a hint about which canonical pages matter on your website. When the URLs inside the sitemap return 200 OK, are indexable, and match their canonical tags, Googlebot gets a clean, trustworthy list. When the sitemap is full of redirects, 404s, and noindex pages, the signal becomes noisy and less useful.
This lesson covers sitemap URLs, sitemap index files, lastmod dates, common sitemap errors, and non-indexable URLs. It also explains how to check sitemap consistency during a technical SEO audit and use SiteAuditLint to find sitemap URLs that do not behave like clean indexable pages.
1. What is an XML sitemap?
An XML sitemap follows the sitemaps.org protocol. A basic sitemap looks like this:
<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<url>
<loc>https://example.com/</loc>
</url>
<url>
<loc>https://example.com/services/</loc>
</url>
<url>
<loc>https://example.com/blog/</loc>
</url>
</urlset>
Each <url> entry contains a URL you want search engines to discover. The <urlset> element wraps the whole list and declares the sitemap namespace.
What belongs in a sitemap?
A sitemap should generally contain URLs that are:
- Canonical, meaning the preferred version of the page.
- Indexable, with no
noindexdirective. - Successful, returning a
200response. - Important to the site and its visitors.
- Accessible to search engines and not blocked by robots.txt.
Avoid treating the sitemap as a complete list of every URL a crawler can technically find. Parameter URLs, filtered views, internal search results, and staging paths do not belong there. The goal is a clean list of URLs you want discovered and considered for indexing.
How a sitemap supports discovery
/sitemap.xml.<loc> and any <lastmod> value.Discovery through the sitemap is a hint. Indexing still depends on the page itself.
2. Sitemap URLs
The <loc> element contains the absolute URL the sitemap recommends to search engines:
<url>
<loc>https://example.com/products/widget/</loc>
</url>
The URL should match the canonical version of the page exactly, including protocol, subdomain, and trailing slash. If https://example.com/page redirects to https://example.com/page/, the sitemap should contain the final canonical URL rather than the redirecting one.
Check your sitemap URLs
Every listed URL should:
- Return
200 OK. - Be indexable.
- Use HTTPS when HTTPS is the canonical version.
- Match the preferred canonical URL.
- Not redirect.
- Not be blocked from crawling.
- Not be a duplicate version of another URL.
Example
| Bad sitemap | Better sitemap |
|---|---|
https://example.com/old-page | https://example.com/ |
https://example.com/page?sort=price | https://example.com/products/ |
https://example.com/private/ | https://example.com/products/widget/ |
https://example.com/missing-page |
The first list mixes a redirected URL, a sorting parameter, a private path, and a missing page. The second gives crawlers a clearer set of canonical URLs to discover.
3. Sitemap index files
A single sitemap file can hold up to 50,000 URLs or 50 MB uncompressed. Large websites split their URLs across several files and reference them from a sitemap index file:
<?xml version="1.0" encoding="UTF-8"?>
<sitemapindex xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<sitemap>
<loc>https://example.com/sitemap-pages.xml</loc>
</sitemap>
<sitemap>
<loc>https://example.com/sitemap-products.xml</loc>
</sitemap>
<sitemap>
<loc>https://example.com/sitemap-blog.xml</loc>
</sitemap>
</sitemapindex>
The index points to individual sitemap files, and a site might organize them by content type:
Sitemap index structure
Splitting by content type also makes Google Search Console reporting more useful, because you can see indexing coverage for products separately from blog posts.
What to check
Index accessibility
The sitemap index itself returns 200 and parses as valid XML.
Referenced files exist
Every child sitemap listed in the index exists and returns a successful response.
Valid child URLs
URLs inside each child sitemap are canonical, indexable, and live.
Retired files removed
Old sitemap files are removed from the index when they are no longer used.
A sitemap index containing broken sitemap references creates unnecessary discovery problems.
4. lastmod dates
The optional <lastmod> element tells search engines when a URL was last significantly modified, using W3C Datetime format:
<url>
<loc>https://example.com/blog/seo-basics/</loc>
<lastmod>2026-09-30</lastmod>
</url>
The date should represent a meaningful change to the page, such as updated main content, structured data, or links, rather than the date the sitemap file was generated. Google uses lastmod when it is consistently and verifiably accurate, and ignores it when it is not. The <changefreq> and <priority> elements are ignored by Google, so accurate lastmod values carry far more weight than either.
Stamping today's date on every URL each time the sitemap regenerates, without changing the content, trains search engines to disregard your lastmod values entirely.
Check your dates
- Missing dates where meaningful update dates are available.
- Invalid date formats.
- Future dates.
- Dates that change without corresponding content changes.
- The same update date applied automatically to every URL.
Accurate lastmod values help search engines prioritize recrawling the URLs that have actually changed.
5. Sitemap errors
Sitemap errors make the file less useful as a discovery source. In Google Search Console, they often surface in the Sitemaps report as statuses like "Couldn't fetch" or "Sitemap could not be read".
404 sitemap
https://example.com/sitemap.xml returns 404 Not Found, so search engines cannot retrieve the sitemap at all.
Invalid XML
Malformed XML, unescaped characters such as a raw &, or incomplete tags can prevent the sitemap from being processed.
Broken sitemap references
An index lists https://example.com/sitemap-products.xml when that file no longer exists.
Incorrect URLs
The sitemap is reachable, but the URLs inside it redirect, return 404 or 5xx, are blocked, carry noindex, or canonicalize elsewhere.
Here is an invalid XML example where the closing tag is incomplete:
<url>
<loc>https://example.com/page</loc>
</url
The sitemap file itself may be accessible while the URLs inside it are problematic, so a successful fetch is only the first check.
6. Non-indexable URLs
One of the most important sitemap checks is whether the listed URLs are actually eligible for indexing. Consider a sitemap that lists https://example.com/product-a/ while the page contains:
<meta name="robots" content="noindex">
The sitemap is recommending a URL that the page itself says should not be indexed. Similarly, if the sitemap lists https://example.com/old-page/ and that URL returns 301 to https://example.com/new-page/, the sitemap should be updated to reference the canonical destination.
Common non-indexable sitemap URLs
| Problem | What it means |
|---|---|
noindex | Page explicitly excluded from indexing |
| Redirect | URL redirects to another URL |
404 | Page no longer exists |
410 | Page has been permanently removed |
| Canonical mismatch | Page points to another canonical URL |
| Blocked URL | Robots.txt prevents crawlers from accessing the URL |
| Server error | URL returns a 5xx response |
For more on how these responses behave, see Lesson 04: Status codes and Lesson 02: Canonicals.
7. Sitemap consistency
A sitemap works best when its signals agree with the rest of the website. Compare two setups for https://example.com/services/seo/:
| Signal | Consistent setup | Conflicting setup |
|---|---|---|
| Sitemap | /services/seo/ | /services/seo/ |
| Canonical | /services/seo/ | /services/ |
| HTTP response | 200 OK | 200 OK |
| Robots meta | index | noindex |
In the conflicting setup, the sitemap says one thing while the canonical tag and the indexing directive say something else. Search engines have to resolve the contradiction themselves, and the sitemap's recommendation loses credibility. When the sitemap, canonical, status code, robots meta, and internal links all point to the same URL, the signal is at its strongest.
8. How to audit sitemaps with SiteAuditLint
A website crawler such as SiteAuditLint helps you identify URLs that appear in your sitemap but do not behave like clean indexable pages.
SiteAuditLint sitemap workflow
- Find the sitemapLocate
/sitemap.xmlor the sitemap referenced by theSitemap:line in your site's robots.txt. - Crawl the siteRun a crawl with SiteAuditLint so sitemap URLs can be compared with what the crawler finds.
- Check sitemap URLsReview URLs discovered through the sitemap and compare their HTTP responses, indexability, canonical URLs, and redirects.
- Find mismatchesFlag sitemap URLs that return errors, redirect, are non-indexable, have canonical mismatches, or no longer exist.
- Clean the sitemapRemove URLs that should not be included and replace outdated URLs with their current canonical versions.
- RecheckRun another crawl after the sitemap has been updated, then resubmit it in Google Search Console.
The goal is not to put every known URL into the sitemap. It is to maintain a reliable list of URLs you want search engines to discover. Also compare in the other direction: important indexable pages missing from the sitemap are worth adding.
9. Practical example
Imagine a website has 1,000 URLs in its sitemap. After crawling them, you find:
| Sitemap URLs | Count | Action |
|---|---|---|
| Valid, indexable URLs | 870 | Keep |
| Redirects | 45 | Replace with final URLs |
| 404 errors | 30 | Remove if permanently gone |
noindex URLs | 25 | Remove if intentionally excluded |
| Canonical mismatches | 20 | Use the canonical URL instead |
| 5xx errors | 10 | Investigate the underlying problem |
| Total | 1,000 |
The sitemap technically contains 1,000 URLs, but only 870 are clean indexable targets. The remaining 130 need investigation. After cleanup, the sitemap provides a much clearer discovery signal, and the gap between "submitted" and "indexed" in the Search Console Sitemaps report becomes easier to interpret.
10. Sitemap checklist
Before considering an XML sitemap clean, work through this list.
XML sitemap audit checklist
0 of 15 tasks completed
11. Key takeaways
- An XML sitemap is a discovery signal, not an indexing command.
- A useful sitemap contains the canonical, indexable URLs you want search engines to discover.
- Sitemap URLs should match the canonical version exactly and return
200 OK. - Sitemap index files should reference only existing, current sitemap files.
lastmoddates should reflect meaningful content changes, and Google ignoreschangefreqandpriority.- Sitemap errors such as a 404 sitemap or invalid XML prevent reliable processing.
- Non-indexable URLs such as redirects, 404s,
noindexpages, and canonical mismatches should not be recommended for discovery. - A clean sitemap keeps discovery signals consistent with canonicals, status codes, and robots directives.
Knowledge check
Ready to go further? Review the technical SEO checklist, or revisit Lesson 04: Status codes.