Independent SiteAuditLint analysis
SiteAuditLint reached 100 URLs on nytimes.com, but the server answered 93 of them, including the homepage, with HTTP 403. Only 6 pages returned content that could be analyzed, so most on-page results in this audit describe a small set of crossword pages.
| Metric | Result |
|---|---|
| Pages crawled | 100 |
| Crawl time | 21s |
| Health score | 85 / 100 (Good) |
| Issue types detected | 16 |
| Critical / warnings / opportunities / info | 99 / 94 / 28 / 2 |
| Audit date | October 4, 2026 |
What stood out
Overall assessment
The dominant observation is access, not on-page quality. The server returned HTTP 403 for 93 of the 100 URLs, and every one of those URLs is a current article or the homepage. A 403 on live editorial content typically reflects access controls applied to automated clients rather than removed pages, so these results describe how nytimes.com responds to a third-party crawler, not how it responds to readers or to search engine crawlers.
That pattern explains the low Technical score of 55 and the 99 critical items. It also means the six pages that did load, all crossword and games pages, carry most of the on-page data. Those pages scored well: titles, meta descriptions, H1s, canonicals, structured data and Open Graph tags were all present.
The clearest deliberate configuration is in robots.txt, which disallows all nine AI crawlers SiteAuditLint checks. That is a policy choice rather than an error, and it is worth reviewing alongside the site's AI search visibility goals.
Audit methodology
SiteAuditLint started from the website's homepage and crawled up to 100 URLs by following discoverable internal links and the URLs listed in the XML sitemap. Results represent the URLs reached during the crawl and do not constitute a complete audit of the entire website.
| Starting URL | https://nytimes.com/ |
|---|---|
| Crawl limit | 100 URLs |
| URLs crawled | 100 (99 HTML) |
| Crawl date | October 4, 2026, 18:05 UTC |
| Crawl duration | 21s |
| Crawl method | Homepage start, internal link and sitemap discovery, robots.txt respected, raw HTML analysis |
| Tool | SiteAuditLint Desktop SEO Crawler |
| Scope | Publicly accessible URLs reached during the crawl |
Technical SEO scorecard
Scores are the category scores SiteAuditLint calculates as part of its Site Health Score.
| Area | Score | Key finding |
|---|---|---|
| Overall health | 85 | Good rating from SiteAuditLint |
| Technical | 55 | 93 URLs returned HTTP 403 to the crawler |
| Content | 100 | No thin content on the 6 analyzable pages |
| Performance | 97 | 1 URL responded in over 1 second (1,035 ms) |
| AEO | 96 | robots.txt blocks all 9 AI crawlers checked |
Major findings
1. URLs returning HTTP 403
HighFinding: SiteAuditLint received HTTP 403 Forbidden for 93 URLs, including the homepage and 92 article URLs from the XML sitemap.
Why it matters: A 403 tells any client it is not allowed to fetch the page. When the same response is sent to search engine or AI crawlers, those pages cannot be crawled. Because the homepage is included, this pattern points to request filtering rather than missing content. See status codes.
| URL | Result |
|---|---|
https://www.nytimes.com/ | HTTP 403 |
https://www.nytimes.com/2026/10/02/briefing/new-york-governor-cornell-jobs-report.html | HTTP 403 |
https://www.nytimes.com/2026/10/03/learning/on-this-day-oct-3.html | HTTP 403 |
Recommended fix: Confirm in server or CDN logs that verified search engine crawlers receive 200 responses for these URLs. If the 403 is limited to unverified bots, no change is needed for search; if not, adjust the bot management rules.
2. Sitemap URLs not returning 200
MediumFinding: 92 of the 98 URLs found in the XML sitemap returned a non-200 status (HTTP 403) to the crawler.
Why it matters: An XML sitemap should list URLs that return 200 and are meant to be indexed. Non-200 entries reduce the reliability of the sitemap as a crawl signal.
| URL | Result |
|---|---|
https://www.nytimes.com/2026/10/02/business/trump-white-house-ban-cnn-politico-ms-now.html | HTTP 403 |
https://www.nytimes.com/2026/10/02/climate/california-trump-administration-fuel-economy.html | HTTP 403 |
Recommended fix: This finding shares a cause with the 403 responses above. Once access is confirmed for verified crawlers, the sitemap entries need no change.
3. AI crawlers blocked in robots.txt
MediumFinding: robots.txt disallows GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, PerplexityBot, Google-Extended, CCBot and Applebot-Extended. No llms.txt file was found.
Why it matters: Blocking these user agents prevents the associated AI products from fetching pages. Some of them, such as OAI-SearchBot and Claude-SearchBot, are used for search and citation rather than training. See robots.txt.
Recommended fix: Treat this as a policy decision: separate training crawlers from search and citation crawlers, and allow any that match the publisher's licensing position.
4. Internal links to a page returning 403
MediumFinding: All 6 analyzable crossword pages link to the homepage, which returned HTTP 403 to the crawler.
Why it matters: SiteAuditLint flags these as links to a broken internal page. Here the target is the homepage, so the flag is a consequence of the 403 response rather than a bad link.
| URL | Evidence |
|---|---|
https://www.nytimes.com/2026/10/02/crosswords/daily-puzzle-2026-10-03.html | Links to https://www.nytimes.com/ [403] |
https://www.nytimes.com/2026/10/03/crosswords/connections-companion-1211.html | Links to https://www.nytimes.com/ [403] |
Recommended fix: No link change is required. Resolve the access behavior described in finding 1.
5. Pages reachable only through the sitemap
LowFinding: The 6 analyzable crossword pages had no internal links pointing to them among the crawled URLs.
Why it matters: Pages without internal links depend on the sitemap for discovery. Because most of the site returned 403, the crawler could not see the pages that would normally link to them, so this result is limited by the crawl. See internal link analysis.
| URL | Evidence |
|---|---|
https://www.nytimes.com/2026/10/03/crosswords/wordle-review-1933.html | No inlinks found |
https://www.nytimes.com/2026/10/03/crosswords/strands-sidekick-945.html | No inlinks found |
Recommended fix: No action is indicated from this crawl alone. Recheck once the crawler can reach section and hub pages.
6. Short and long titles, shared descriptions
LowFinding: 3 titles are over 60 characters, 1 title is under 30 characters, 2 pages share a meta description, and 1 meta description is under 70 characters.
Why it matters: Titles and descriptions are the main inputs for search snippets. Shared descriptions make two pages look alike in results. See on-page SEO checks.
| URL | Result |
|---|---|
https://www.nytimes.com/2026/10/03/crosswords/spelling-bee-forum.html | Title 65 characters, description 45 characters |
https://www.nytimes.com/2026/10/03/crosswords/daily-puzzle-2026-10-04.html | Title 25 characters |
https://www.nytimes.com/2026/10/03/crosswords/strands-sidekick-945.html | Description shared with connections-companion-1211 |
Recommended fix: Adjust the crossword page templates so titles fall between 30 and 60 characters and each puzzle page generates a distinct description.
What's working well
- All 6 analyzable pages have a single H1, a self-referencing canonical and a meta description.
- Structured data was present on every analyzable page, with no invalid JSON-LD.
- No mixed content, missing HSTS, missing viewport or HTTP pages were detected.
- The non-www domain 301 redirects to https://www.nytimes.com/ in a single hop.
- No 5xx server errors and only 1 response over 1 second across 100 URLs.
Priority fixes
- Verify crawler access for search enginesCheck that verified search engine crawlers receive 200 responses where SiteAuditLint received 403. Pages affected: 93
- Review the AI crawler policy in robots.txtDecide whether search and citation agents should be treated differently from training agents. Pages affected: 1
- Confirm sitemap entries resolve for verified crawlersThe 92 non-200 sitemap entries share the 403 cause. Pages affected: 92
- Tune crossword page titles and descriptionsTemplate-level changes cover the length and duplicate findings. Pages affected: 6
Crawl data
| Metric | Result |
|---|---|
| URLs crawled | 100 |
| HTML pages | 99 |
| 2xx responses | 6 |
| 3xx redirects | 1 |
| 4xx responses (all 403) | 93 |
| 5xx responses | 0 |
| URLs in XML sitemap | 98 |
| Sitemap URLs not returning 200 | 92 |
| Missing titles | 0 |
| Missing meta descriptions | 0 |
| Duplicate meta descriptions | 2 |
| Missing H1s | 0 |
| Self-referencing canonicals | 6 of 6 |
| Slow responses (over 1,000 ms) | 1 |
| AI crawlers allowed | 0 of 9 |
On-page counts (titles, descriptions, headings, canonicals) are measured across the 6 URLs that returned an HTML page with a 200 status.
Audit limitations
The crawl was limited to 100 URLs and started from the homepage, so only URLs discovered through internal links and the XML sitemap are represented. A 100-URL crawl does not represent the entire website. Findings reflect the site at the time of the crawl on October 4, 2026. SiteAuditLint analyzed the HTML returned by the server, so content added later by JavaScript and responses that differ for automated clients may not match what a browser shows. This audit evaluates technical observations, not overall business or search performance.
Because 93 URLs returned 403, on-page metrics describe only 6 pages and should not be read as representative of nytimes.com as a whole.
This is an independent SiteAuditLint analysis of publicly accessible pages. It was not requested or endorsed by The New York Times, and it does not imply that The New York Times uses SiteAuditLint.