Home›Academy›Site Audits›Audit process

Site Audits · Lesson 01 · Workflow

Audit process

How to plan and run a site audit, from defining scope and configuring the crawler to collecting supporting data and saving a baseline for future comparison.

04 Site Audits11 min Read timeIntermediate Level

01 / Workflow

A site audit is a structured review of how a website is crawled, rendered, indexed, and understood. The audit process turns a crawl of thousands of URLs into a short list of evidence-backed problems, so the work starts with scope and configuration rather than with a list of warnings.

Earlier courses explained individual signals such as status codes, canonical tags, and title tags. This lesson shows how to combine them into a repeatable audit workflow: defining the audit scope, configuring the crawler, collecting supporting data, establishing a baseline, and planning a recrawl after fixes ship.

The same process works for a pre-launch review, a site migration, a sudden traffic drop, or a recurring monthly health check. Only the scope and the questions change.

1. What a site audit is

A technical SEO audit answers three questions about a website: can search engines and AI crawlers reach the important pages, can they index and render them correctly, and do those pages clearly describe their topic? Every check in an audit supports one of those questions.

Audit layerCore questionTypical signals
CrawlabilityCan bots reach the URL?robots.txt rules, internal links, HTTP status codes, crawl depth
IndexabilityIs the URL allowed and chosen for the index?noindex, canonical URL, duplicate content, XML sitemap inclusion
RenderingDoes the content exist after JavaScript runs?Rendered HTML, JavaScript links, lazy-loaded content
Page contentDoes the page clearly describe its topic?Title tags, meta descriptions, headings, thin content
Site architectureHow do pages connect?Internal linking, orphan pages, redirect chains, click depth
An audit is a diagnosis, not a score

A site health score summarizes the crawl, but it does not tell you why traffic changed or which fix matters most. Treat the score as a trend line and the individual findings as the evidence.

2. Define the audit scope and goal

Before crawling, write down why the audit is happening. The goal decides which sections, subdomains, and signals deserve the most attention.

Audit triggerScopeFocus
Pre-launch or redesignStaging environment plus the live site for comparisonNoindex left on templates, blocked robots.txt, missing redirects
Site migrationOld URL list mapped against the new site301 redirects, redirect chains, canonical changes, lost pages
Traffic or ranking dropAffected sections or templatesIndexability changes, status code spikes, content changes
Recurring health checkFull site on a fixed scheduleNew issues since the previous crawl
AI visibility reviewKey content and commercial pagesAI crawler access, rendered content, structured data

Scope also covers boundaries: which subdomains are included, whether URL parameters should be crawled, and whether paginated or faceted URLs belong in the audit. For a pre-launch audit, the pre-launch website checklist gives a practical starting list.

3. Configure the crawl

Crawl configuration decides what the crawler sees. A crawl with the wrong settings produces findings that do not match what Googlebot or an AI crawler experiences.

SettingWhy it matters
Start URLUse the canonical protocol and hostname, for example the HTTPS www version.
User agentCrawling as a search engine bot can reveal rules that treat bots differently from browsers.
Respect robots.txtShows what compliant bots can reach. Disable only to inspect blocked sections deliberately.
JavaScript renderingRequired when navigation, content, or links are generated client-side.
Crawl speed and limitsProtects the server and avoids HTTP 429 rate limiting.
URL parametersPrevents infinite crawl spaces from filters, sorting, and session IDs.
Crawl depthLimits how many clicks from the start URL the crawler follows.

The SiteAuditLint crawl settings guide explains each option, and the JavaScript rendering guide covers when to enable rendering. If a crawl slows or stops early, check the rate limiting guide before assuming the site is broken.

Staging needs authentication, not guesses

Staging sites are often password protected or blocked in robots.txt. Configure credentials and crawler rules for staging explicitly, and remember that settings such as a sitewide noindex must be removed before launch.

4. Collect supporting data

A crawl shows what the site exposes. It does not show what search engines actually chose to do. Strong audits combine the crawl with other sources:

Google Search Console

Page indexing reports, crawl stats, and which URLs Google selected as canonical.

XML sitemaps

The list of URLs the site owner says are important, compared against the crawl.

Server log files

Which URLs Googlebot, Bingbot, and AI crawlers really requested, and how often.

Analytics

Which pages earn traffic and conversions, so fixes can be weighted by value.

Backlink data

Which URLs have external links, so broken or redirected linked pages get priority.

Comparing sources reveals gaps. A URL in the sitemap that the crawler never found through internal links is a likely orphan page. A URL that earns backlinks but returns a 404 is losing value. These cross-source findings are often the most valuable in an audit.

5. Run the crawl and check completeness

After the crawl finishes, confirm it actually represents the site before reading any issues.

  1. Compare URL countsCheck the number of crawled HTML pages against the sitemap and the indexed page count in Search Console.
  2. Look for crawl stopsA large share of 403, 429, or 5xx responses often means the crawler was blocked or throttled.
  3. Check discovered sectionsConfirm that every major folder or template appears in the crawl.
  4. Review rendered contentOn JavaScript sites, verify that key text and links appear in the rendered HTML.
  5. Note what was excludedRecord parameters, subdomains, or folders left out so findings are read in context.

Spotting an incomplete crawl

SuspiciousSitemap: 4,200 URLs
Crawled: 310 URLs
Most responses after URL 300 are HTTP 403.
ReliableSitemap: 4,200 URLs
Crawled: 4,160 URLs
The crawler was allowlisted and speed reduced.

A firewall or bot protection rule often blocks audit crawlers. The same rule may also block AI crawlers.

If blocking is the cause, the guide on fixing 403 errors that block AI crawlers walks through the usual firewall and CDN settings.

6. Establish a baseline

The first complete crawl becomes the baseline. Save it with the date, the configuration, and a short note about the site state, such as a recent release or migration. Without a baseline, you cannot prove that a fix worked or that a new release introduced a regression.

Useful baseline metrics include indexable URL count, non-200 URL count, redirect count, pages with missing titles or H1s, orphan pages, and average crawl depth of important templates.

7. Recrawl and audit cadence

An audit is a loop, not a single event. After fixes ship, recrawl with the same configuration and compare the results against the baseline.

Site typeSuggested cadence
Small brochure siteQuarterly, plus after any redesign
Blog or content siteMonthly
Ecommerce or large catalogWeekly or after every major release
During a migrationBefore launch, on launch day, and during the following weeks

SiteAuditLint scheduled audits can run the same crawl automatically, and the compare view shows which issues are new, fixed, or unchanged between crawls.

8. Running the audit process in SiteAuditLint

  1. Set the start URL and scopeUse the canonical hostname and decide which subdomains and parameters to include.
  2. Configure crawl settingsChoose the user agent, rendering, speed, and robots.txt behavior.
  3. Run the first auditFollow the first audit guide and let the crawl complete.
  4. Validate completenessCompare the URL count with the sitemap and check for blocked or throttled responses.
  5. Save the baselineKeep the crawl for later comparison and schedule the next run.

9. Practical exercise: Plan an audit

Pick a website you manage and write a one-page audit plan before crawling. Include the goal, scope, crawl settings, data sources, and the date of the next recrawl.

Audit process checklist

0 of 10 tasks completed

Key takeaways

  • A site audit reviews crawlability, indexability, rendering, page content, and site architecture.
  • The audit goal defines the scope and the checks that matter most.
  • Crawl configuration must match how search engines and AI crawlers experience the site.
  • Search Console, sitemaps, logs, analytics, and backlink data add evidence the crawl cannot provide.
  • Always confirm a crawl is complete before trusting its findings.
  • A saved baseline and a scheduled recrawl turn a one-off audit into a repeatable process.

Knowledge check

1. What should happen before you start a crawl?
2. A crawl finds 300 URLs but the sitemap lists 4,000. Most late responses are 403. What is the likely cause?
3. Why save a baseline crawl?
4. Which source shows what Googlebot actually requested?

Ready to go further? Read the technical SEO audit guide, or revisit how bots reach pages in the Crawling lesson.