Home›Academy›Site Audits›Issue checks

Site Audits · Lesson 02 · Detection

Issue checks

What each category of audit check measures, how crawlers detect problems across pages, and how to verify flagged issues before reporting them.

04 Site Audits12 min Read timeIntermediate Level

02 / Detection

Issue checks are the individual tests a crawler runs against every URL. Each check compares a page against a rule, such as "returns HTTP 200" or "has one descriptive title tag", and flags URLs that fail. Knowing what each check measures, and what it cannot measure, is what separates a useful audit from a long export of warnings.

This lesson groups the most important checks into categories: HTTP responses and redirects, indexability, canonicalization and duplication, on-page elements, internal linking, security, structured data, and AI crawler access. It also explains severity levels, false positives, and how to verify a flagged issue before reporting it.

1. How issue checks work

A crawler fetches each URL, records the response headers and HTML, optionally renders JavaScript, and then evaluates rules. A check can look at a single page, such as a missing meta description, or at relationships between pages, such as two URLs sharing the same title or a page with no incoming internal links.

Check typeEvaluatesExample
Page-levelOne URL in isolationMissing H1, title too long, missing lang attribute
Response-levelHTTP status and headers4xx, 5xx, missing HSTS header, slow response
Cross-pageRelationships between URLsDuplicate titles, near-duplicate content, orphan pages
Link-levelSource and target of each linkLinks to broken pages, links to redirects, empty anchors
Site-levelFiles that apply to the whole siteMissing robots.txt, missing XML sitemap, AI crawlers blocked

2. HTTP status and redirect checks

Response checks come first because a page that does not return HTTP 200 cannot rank, regardless of how well it is optimized.

4xx client errors

Broken URLs such as 404 and 410. See the 4xx issue guide.

5xx server errors

The server failed to respond. Often intermittent and worth rechecking.

Redirect chains

A redirect points to another redirect before the final URL.

Redirect loops

Redirects that never resolve to a final page.

Internal links to redirects

Links that should point directly at the final URL.

Slow response

Time to first byte is high enough to slow crawling.

The Redirects lesson explains when to use 301 versus 302 redirects, and the broken link audit guide shows how to trace broken links back to the pages that contain them.

3. Indexability checks

Indexability checks confirm whether a page is eligible to appear in search results.

CheckWhat it flagsVerify by
NoindexMeta robots or X-Robots-Tag set to noindexConfirming whether the page should be excluded
Blocked by robots.txtA disallow rule prevents crawlingTesting the rule against the URL path
CanonicalisedThe canonical points to a different URLChecking the target is the preferred version
Sitemap contains noindex URLsConflicting signals between sitemap and pageRemoving noindex URLs from the sitemap
Sitemap contains non-200 URLsRedirected or broken URLs listed as importantUpdating the sitemap to final URLs
Conflicting signals are the real problem

A single noindex is often intentional. The audit value comes from conflicts: a noindex page listed in the XML sitemap, a canonical pointing to a redirected URL, or an important page blocked in robots.txt while receiving internal links.

For fixes, see noindex issues, canonical tag issues, and XML sitemap errors.

4. Duplicate and thin content checks

Duplication checks compare pages against each other. Exact duplicates share identical content, while near-duplicates differ only slightly, such as product variants or location pages built from one template. Thin content checks flag pages with very little unique main content.

Common causes include URL parameters, HTTP and HTTPS versions, trailing slash variations, uppercase URLs, and printer-friendly pages. The duplicate content guide and the thin content guide cover consolidation options.

5. On-page element checks

ElementCommon checks
Title tagMissing, duplicate, too long, too short, multiple title elements
Meta descriptionMissing, duplicate, too long, too short
HeadingsMissing H1, multiple H1s, missing H2s, skipped levels
ImagesMissing alt text on meaningful images
HTML documentMissing lang attribute, missing viewport, large HTML size
Social metadataMissing Open Graph tags

Length checks are guidance, not rules. A title flagged as long may still work if the important words appear first. Review the Titles, Meta descriptions, and Headings lessons for the reasoning behind each threshold.

Orphan pages

Pages found in the sitemap but not linked internally.

Deep pages

Important pages many clicks from the homepage.

Empty anchors

Links with no accessible text.

Generic anchors

Anchor text such as click here or read more.

HTTP links on HTTPS pages

Internal links that force an insecure request or redirect.

Links to broken external URLs

Outbound links that return errors.

Architecture checks connect directly to the Internal links lesson. Related fixes are covered in fixing broken internal links and HTTPS pages linking to HTTP.

7. Security, structured data, and AI crawler checks

AreaChecks
SecurityNot HTTPS, mixed content, missing HSTS header
Structured dataMissing schema markup, invalid schema properties
AI crawler accessGPTBot, ClaudeBot, PerplexityBot, and Google-Extended blocked by robots.txt or firewall rules
AI discovery filesMissing llms.txt, where the site has decided to publish one
Answer readinessPages with no question-style headings or direct answers

AI crawler checks matter because a site can rank in traditional search while being invisible to AI answer engines. See the robots.txt and AI crawlers guide and the article on checking crawlability for Google, Bing, and AI bots.

8. Severity levels and false positives

Crawlers assign a severity to each check, typically error, warning, or notice. Severity describes the potential impact of the rule, not the impact on your specific site.

Reading a flag in context

FlaggedNoindex on /cart/
Severity: warning
The check correctly detects noindex.
DecisionIntentional
No action needed
Cart pages should not be indexed.

The check is accurate. Whether it is a problem depends on the page's purpose.

  1. Open the URLConfirm the issue exists in the live page, not only in the crawl data.
  2. Check the rendered versionSome elements only appear after JavaScript runs.
  3. Compare with Search ConsoleSee whether Google reports the same condition.
  4. Confirm intentAsk whether the condition is deliberate, such as noindex on internal search pages.
  5. Record the decisionMark accepted exceptions so they do not reappear in every report.

The severity levels guide explains how SiteAuditLint classifies each check.

9. Practical exercise: Verify ten flagged issues

From your latest crawl, pick ten flagged URLs across different categories. For each one, confirm the issue manually and decide whether it is a real problem, an accepted exception, or a false positive.

Issue check verification

0 of 8 tasks completed

Key takeaways

  • Issue checks test pages, responses, links, relationships, and site-wide files.
  • Status code and indexability checks come first because they decide whether a page can rank.
  • Conflicting signals are usually more important than single flags.
  • Length thresholds for titles and descriptions are guidance, not strict rules.
  • AI crawler access is now part of a complete issue check.
  • Severity describes potential impact. Verify each flag before reporting it.

Knowledge check

1. Which issue is the strongest sign of a real problem?
2. What is a cross-page check?
3. A crawler flags a warning. What should you do first?
4. Why include AI crawler checks in an audit?

Continue with the SEO, GEO, and AEO audit checklist, or browse the full list of SiteAuditLint issue definitions.