Crawl settings explained
What each crawl setting does, from maximum pages, parallel requests and delay to click depth, subdomains, sitemaps, robots.txt and the user agent.
Open Settings > Crawl Settings... to control how SiteAud!tLint crawls. Scheduled audits use the same saved settings.
| Setting | Default | What it does |
|---|---|---|
| Maximum pages | Edition maximum | Stops the crawl after this many pages. 0 follows your edition, so upgrading lifts the limit without changing settings. |
| Parallel requests | 5 | How many pages are fetched at the same time (1 to 20). |
| Delay between requests | 0 s | Pause each worker takes between requests, in steps of 0.25 seconds. |
| Request timeout | 20 s | How long to wait for a page before recording it as a fetch error. |
| Maximum click depth | Unlimited | How many clicks from the start URL to follow. |
| Respect robots.txt | On | Skips URLs your robots.txt disallows and reports them as blocked. |
| Include subdomains | Off | Treats blog.example.com and shop.example.com as part of the site. |
| Discover and crawl XML sitemaps | On | Reads sitemaps from robots.txt and /sitemap.xml to find pages that links miss. |
| Check external links | On | Checks the status code of every outbound link after the crawl. |
| Render JavaScript (Chromium) | Off | Loads pages in a real browser. Only shown when the optional component is installed. |
| User agent | SiteAuditLintBot | The name the crawler sends with each request. |
| Custom search | None | Up to 10 text or regex rules run against page source. |
Speed versus completeness
More parallel requests finish faster, but some hosts answer fast crawls with HTTP 429 "Too Many Requests". SiteAud!tLint slows down automatically when that happens. For Shopify stores and shared hosting, 1 to 2 parallel requests with a 1 second delay gives the most complete audit. See Rate limiting and HTTP 429.
Crawling a staging site
Staging sites are often blocked in robots.txt or protected by a login. To audit one, untick Respect robots.txt for that audit only. Pages behind a login screen cannot be crawled.
Firewalls and bot protection
If your host or CDN blocks unknown bots, pages come back as Page could not be fetched or 403 errors. Allow the SiteAuditLintBot user agent in your firewall, or change User agent to one your firewall allows. See About SiteAuditLintBot.
Custom source code search
Custom search finds pages whose HTML contains, or does not contain, a string or regular expression. Common uses: pages still loading an old script, pages missing a required disclaimer, or pages using a deprecated CSS class. See Custom source code search.
Applies to SiteAud!tLint 0.3 and later. Current version 0.7.1.