02 / Detection
Issue checks are the individual tests a crawler runs against every URL. Each check compares a page against a rule, such as "returns HTTP 200" or "has one descriptive title tag", and flags URLs that fail. Knowing what each check measures, and what it cannot measure, is what separates a useful audit from a long export of warnings.
This lesson groups the most important checks into categories: HTTP responses and redirects, indexability, canonicalization and duplication, on-page elements, internal linking, security, structured data, and AI crawler access. It also explains severity levels, false positives, and how to verify a flagged issue before reporting it.
1. How issue checks work
A crawler fetches each URL, records the response headers and HTML, optionally renders JavaScript, and then evaluates rules. A check can look at a single page, such as a missing meta description, or at relationships between pages, such as two URLs sharing the same title or a page with no incoming internal links.
| Check type | Evaluates | Example |
|---|---|---|
| Page-level | One URL in isolation | Missing H1, title too long, missing lang attribute |
| Response-level | HTTP status and headers | 4xx, 5xx, missing HSTS header, slow response |
| Cross-page | Relationships between URLs | Duplicate titles, near-duplicate content, orphan pages |
| Link-level | Source and target of each link | Links to broken pages, links to redirects, empty anchors |
| Site-level | Files that apply to the whole site | Missing robots.txt, missing XML sitemap, AI crawlers blocked |
2. HTTP status and redirect checks
Response checks come first because a page that does not return HTTP 200 cannot rank, regardless of how well it is optimized.
4xx client errors
Broken URLs such as 404 and 410. See the 4xx issue guide.
5xx server errors
The server failed to respond. Often intermittent and worth rechecking.
Redirect chains
A redirect points to another redirect before the final URL.
Redirect loops
Redirects that never resolve to a final page.
Internal links to redirects
Links that should point directly at the final URL.
Slow response
Time to first byte is high enough to slow crawling.
The Redirects lesson explains when to use 301 versus 302 redirects, and the broken link audit guide shows how to trace broken links back to the pages that contain them.
3. Indexability checks
Indexability checks confirm whether a page is eligible to appear in search results.
| Check | What it flags | Verify by |
|---|---|---|
| Noindex | Meta robots or X-Robots-Tag set to noindex | Confirming whether the page should be excluded |
| Blocked by robots.txt | A disallow rule prevents crawling | Testing the rule against the URL path |
| Canonicalised | The canonical points to a different URL | Checking the target is the preferred version |
| Sitemap contains noindex URLs | Conflicting signals between sitemap and page | Removing noindex URLs from the sitemap |
| Sitemap contains non-200 URLs | Redirected or broken URLs listed as important | Updating the sitemap to final URLs |
A single noindex is often intentional. The audit value comes from conflicts: a noindex page listed in the XML sitemap, a canonical pointing to a redirected URL, or an important page blocked in robots.txt while receiving internal links.
For fixes, see noindex issues, canonical tag issues, and XML sitemap errors.
4. Duplicate and thin content checks
Duplication checks compare pages against each other. Exact duplicates share identical content, while near-duplicates differ only slightly, such as product variants or location pages built from one template. Thin content checks flag pages with very little unique main content.
Common causes include URL parameters, HTTP and HTTPS versions, trailing slash variations, uppercase URLs, and printer-friendly pages. The duplicate content guide and the thin content guide cover consolidation options.
5. On-page element checks
| Element | Common checks |
|---|---|
| Title tag | Missing, duplicate, too long, too short, multiple title elements |
| Meta description | Missing, duplicate, too long, too short |
| Headings | Missing H1, multiple H1s, missing H2s, skipped levels |
| Images | Missing alt text on meaningful images |
| HTML document | Missing lang attribute, missing viewport, large HTML size |
| Social metadata | Missing Open Graph tags |
Length checks are guidance, not rules. A title flagged as long may still work if the important words appear first. Review the Titles, Meta descriptions, and Headings lessons for the reasoning behind each threshold.
6. Internal link and architecture checks
Orphan pages
Pages found in the sitemap but not linked internally.
Deep pages
Important pages many clicks from the homepage.
Empty anchors
Links with no accessible text.
Generic anchors
Anchor text such as click here or read more.
HTTP links on HTTPS pages
Internal links that force an insecure request or redirect.
Links to broken external URLs
Outbound links that return errors.
Architecture checks connect directly to the Internal links lesson. Related fixes are covered in fixing broken internal links and HTTPS pages linking to HTTP.
7. Security, structured data, and AI crawler checks
| Area | Checks |
|---|---|
| Security | Not HTTPS, mixed content, missing HSTS header |
| Structured data | Missing schema markup, invalid schema properties |
| AI crawler access | GPTBot, ClaudeBot, PerplexityBot, and Google-Extended blocked by robots.txt or firewall rules |
| AI discovery files | Missing llms.txt, where the site has decided to publish one |
| Answer readiness | Pages with no question-style headings or direct answers |
AI crawler checks matter because a site can rank in traditional search while being invisible to AI answer engines. See the robots.txt and AI crawlers guide and the article on checking crawlability for Google, Bing, and AI bots.
8. Severity levels and false positives
Crawlers assign a severity to each check, typically error, warning, or notice. Severity describes the potential impact of the rule, not the impact on your specific site.
Reading a flag in context
Severity: warningThe check correctly detects noindex.
No action neededCart pages should not be indexed.
The check is accurate. Whether it is a problem depends on the page's purpose.
- Open the URLConfirm the issue exists in the live page, not only in the crawl data.
- Check the rendered versionSome elements only appear after JavaScript runs.
- Compare with Search ConsoleSee whether Google reports the same condition.
- Confirm intentAsk whether the condition is deliberate, such as noindex on internal search pages.
- Record the decisionMark accepted exceptions so they do not reappear in every report.
The severity levels guide explains how SiteAuditLint classifies each check.
9. Practical exercise: Verify ten flagged issues
From your latest crawl, pick ten flagged URLs across different categories. For each one, confirm the issue manually and decide whether it is a real problem, an accepted exception, or a false positive.
Issue check verification
0 of 8 tasks completed
Key takeaways
- Issue checks test pages, responses, links, relationships, and site-wide files.
- Status code and indexability checks come first because they decide whether a page can rank.
- Conflicting signals are usually more important than single flags.
- Length thresholds for titles and descriptions are guidance, not strict rules.
- AI crawler access is now part of a complete issue check.
- Severity describes potential impact. Verify each flag before reporting it.
Knowledge check
Continue with the SEO, GEO, and AEO audit checklist, or browse the full list of SiteAuditLint issue definitions.