A 403 Forbidden error can prevent AI crawlers from accessing your website, even when your pages are publicly available to human visitors. If you want your content to be discoverable by AI-powered search engines and answer engines, it is important to understand why these errors occur and how to resolve them without weakening your website's security.
In many cases, the problem is not the content itself. It is the security layer protecting the website. Web application firewalls, bot protection, hosting-level restrictions, and server configurations can block automated crawlers before they reach your pages. Learn more about how to check website crawlability for Google, Bing, and AI bots to identify access issues.
The following sections explain how to diagnose and fix 403 errors affecting AI crawlers, with a focus on Cloudflare, shared hosting, and server-side security settings.
1. Why AI crawlers receive 403 errors
A 403 error means the server or a security system understood the request but refused to fulfill it. The response can come from your hosting server, a firewall, a CDN, or an application-level security rule.
From a website owner's perspective, one common cause is that additional security measures have been enabled to protect the website from unwanted automated traffic. These protections may also block legitimate AI crawlers.
For example, a website may:
- Block requests identified as automated. A firewall may treat a crawler as suspicious simply because it does not behave like a typical browser.
- Require browser challenges. Some bot protection systems expect JavaScript execution or a browser verification step that a crawler cannot complete.
- Restrict IP addresses or locations. Network-level rules may reject requests from specific IP ranges, hosting providers, or countries.
- Apply restrictive rate limits. Repeated requests can trigger temporary blocks when the configured thresholds are too low.
- Filter unfamiliar user agents. Security rules may block user-agent strings associated with automated clients.
- Restrict access through server configuration. Rules in the web server, application, or hosting control panel may deny requests before the page is delivered.
The important distinction is that a 403 does not automatically mean your website has disallowed AI crawling intentionally. It may be an unintended result of a security configuration.
2. Identify which AI crawler is being blocked
Before changing any security settings, identify which crawler is receiving the error. Different AI services use different crawlers, and allowing one does not automatically allow the others.
| AI service | Crawler | Main purpose |
|---|---|---|
| OpenAI | OAI-SearchBot |
Discovering content for ChatGPT search. |
| OpenAI | GPTBot |
Crawling content that may be used to train AI models. |
| Anthropic | ClaudeBot |
Crawling content for Anthropic's systems. |
| Perplexity | PerplexityBot |
Discovering and indexing web content for search. |
Googlebot |
Crawling pages for Google Search. |
For OpenAI, OAI-SearchBot and GPTBot serve separate purposes. Allowing the search crawler does not require allowing the training crawler.
To understand how automated bots access your website, see what a website crawler is and how it works. You can also use our article on checking whether ChatGPT or Perplexity cites your site to investigate your visibility in AI search.
Check your server logs to establish:
- The exact user agent: Which crawler identity appeared in the request?
- The affected URL: Does the 403 occur on the homepage, selected pages, or the entire website?
- The request timestamp: When did the failed request occur, including the time zone?
- The response source: Did Cloudflare, your hosting server, or your application return the error?
- The blocking rule: Is there an event ID, firewall rule ID, or error message that explains the denial?
Do not assume that a request is legitimate solely because its user agent claims to be an AI crawler. User-agent strings can be spoofed. Where available, use the crawler operator's published IP ranges or your security provider's verified-bot identification features to validate the traffic.
3. Check your robots.txt file first
Your robots.txt file is one of the first places to check when investigating crawl access. It tells compliant crawlers which parts of your website they may request.
However, robots.txt does not override a firewall or grant access to a server that is returning 403 errors. A crawler can be permitted by robots.txt and still be blocked by another security layer.
For example, if you want to allow OpenAI's search crawler while keeping its training crawler out, you could use:
User-agent: OAI-SearchBot
Allow: /
User-agent: GPTBot
Disallow: /
This example allows the search crawler to request pages across the site while disallowing the training crawler. Add or retain other directives your site needs, and check for conflicting rules elsewhere in your file.
If your goal is to make your website accessible to a particular AI search crawler, check that the crawler is not accidentally disallowed. Then test whether the server actually permits its requests.
A robots.txt file is only one part of crawlability. For a broader review of technical access issues, follow the technical SEO checklist for search and AI.
For Google, a 403 is a client error and is not a recommended way to manage crawl rates. Google advises using appropriate rate-limiting methods, such as 429 Too Many Requests, rather than returning other 4xx errors to slow its crawler.
4. Fix 403 errors caused by Cloudflare
Cloudflare is a common part of the troubleshooting process because it sits between visitors and your origin server. Its security features can block requests before they ever reach your hosting provider.
Cloudflare's WAF, bot protection, custom firewall rules, and AI crawler controls can all affect whether an automated request is permitted. Cloudflare also offers a setting specifically designed to block AI bots, so check that you have not enabled it unintentionally.
For a closer look at monitoring AI bot activity, read Cloudflare for AI monitoring: How to track AI crawlers. The article covers monitoring crawler activity and understanding which automated clients are accessing your website.
Cloudflare troubleshooting flow
Follow the request from the security event to the rule that needs attention.
Step 1: Check Cloudflare Security Events
Start by logging in to your Cloudflare dashboard and selecting the affected website.
- Open Security → Events (the exact dashboard label may vary).
- Filter the events by the time the crawler received the 403.
- Look for requests to the affected URLs.
- Inspect the action, rule, source IP, user agent, and other available event details.
- Identify the rule or protection feature responsible for the block.
If Cloudflare generated the 403, its event information can help identify which rule needs attention. If the error page is not branded by Cloudflare, the origin server may be returning it instead.
Step 2: Review AI crawler settings
Cloudflare provides AI Crawl Control features for managing AI crawlers. Depending on your account and configuration, you can inspect crawler activity and choose to allow or block specific crawlers.
If you want to make your content available to a particular crawler:
- Open your Cloudflare dashboard.
- Select your website.
- Navigate to AI Crawl Control.
- Review the crawler activity and security information.
- Find the crawler you want to permit.
- Set its action to allow, where the feature is available and appropriate for your site's content policy.
Remember that allowing a crawler in AI Crawl Control does not necessarily bypass every other security rule. A separate WAF rule or origin-server restriction may still block the request.
Step 3: Review WAF and bot protection rules
Look for rules that might block legitimate crawlers based on:
- User-agent strings: Conditions that reject specific automated clients or broad groups of browser-like requests.
- Request frequency: Rate limits that are too strict for the crawler's normal request pattern.
- IP address or network: Deny rules that match an address or range used by a legitimate crawler.
- Country or region: Geographic restrictions that affect requests originating outside your expected visitor locations.
- Request paths: Rules that block directories, content types, or URL patterns that the crawler needs to access.
- Bot detection scores: Automated traffic rules that incorrectly classify a legitimate crawler as malicious.
A broad rule that blocks all automated traffic can also affect crawlers you want to allow. Review its conditions and actions, then adjust the rule to preserve the intended protection while avoiding false positives.
If your Cloudflare plan supports custom rules and verified-bot conditions, consider using those capabilities to make exceptions for validated, legitimate crawlers. Do not rely on a user-agent-only allowlist as your sole security measure, because an attacker can impersonate a crawler's user agent.
Step 4: Test after making changes
Once you have adjusted the relevant security setting, test the affected URL again.
- Confirm that the crawler can access the page.
- Check whether the response is
200 OKfor a normal, publicly accessible page. - Verify whether the 403 still occurs on other URLs.
- Make sure your firewall continues to block unrelated suspicious requests.
- Confirm that your
robots.txtfile remains accessible.
Allow time for configuration changes to take effect and examine fresh security events. A successful test with one crawler does not guarantee that other AI crawlers have access.
5. Fix 403 errors on shared hosting
Shared hosting introduces another possible source of 403 errors. On these plans, multiple websites use resources on the same physical server, and the hosting provider may apply security restrictions at the server level.
Even if you have not configured a firewall yourself, your hosting provider may have enabled protections such as ModSecurity, IP restrictions, anti-bot filters, or request limits.
This is especially important when you use a managed hosting service and do not have direct access to the server's configuration.
Common shared hosting causes
| Potential cause | How it can affect AI crawlers | What to check |
|---|---|---|
| ModSecurity | A security rule may flag an automated request as suspicious. | Ask your host to inspect the ModSecurity audit logs and identify the triggered rule ID. |
| IP restrictions | Requests from certain IP ranges may be denied. | Check IP deny rules and hosting firewall settings. |
| User-agent filtering | Rules may reject requests from unfamiliar or automated clients. | Review .htaccess and hosting security settings. |
| Rate limiting | Repeated requests may trigger temporary blocks. | Review request limits, access logs, and any temporary block settings. |
| Directory permissions | Incorrect permissions can prevent access to pages or files. | Check file and directory permissions against the hosting provider's recommended settings. |
| Hotlink or access protection | Restrictions may interfere with requests to particular resources. | Review resource protection settings and the affected URLs. |
These are possible causes, not a diagnosis. Check the server logs or ask your provider to identify the exact rule responsible before making changes. Cloudflare's troubleshooting documentation also identifies origin permissions, ModSecurity, and IP deny rules as common causes of 403 responses.
If the error is part of a wider website accessibility problem, consider running a technical SEO audit. A structured audit can help uncover additional technical problems alongside crawler access restrictions.
Step 1: Inspect your .htaccess file
For websites running on Apache, the .htaccess file can contain access restrictions and rewrite rules that affect incoming requests.
For example, a rule that blocks particular user agents or denies requests from a specified IP range may unintentionally affect legitimate crawlers.
Review any rules that use directives such as:
Require all denied— a rule that can deny access to a directory or resource.Deny from— an access control directive that may be used in older Apache configurations.RewriteCond %{HTTP_USER_AGENT}— a condition that can match user-agent strings.RewriteRule— a rewrite or blocking rule that may affect requests based on their URL or other conditions.- IP-based access restrictions — rules that deny requests from specific addresses or networks.
Do not delete security rules without understanding their purpose. If you find a rule that appears to be responsible, test a narrowly scoped adjustment and confirm that the website remains protected.
Step 2: Contact your hosting provider
If you do not have access to the server's security logs or firewall configuration, ask your hosting provider for help.
Provide the affected URL, the time of the failed request, the crawler's user-agent string, the response code, and any request ID or error details you have.
Subject: Request to investigate 403 errors affecting AI crawlers
Hello Support Team,
I am investigating HTTP 403 Forbidden errors affecting AI crawlers on my website.
The affected URLs are publicly accessible to regular visitors, but requests from AI crawlers appear to be blocked.
Could you please check your server logs, ModSecurity audit logs, firewall rules, and any automated bot protection that may be causing these requests to return 403?
If a security rule is responsible, please let me know which rule is triggering the block and whether a targeted exception can be applied for verified, legitimate crawlers.
I can provide the affected URLs, request timestamps, user-agent strings, and any available error or request IDs.
Thank you.
If the hosting provider confirms that a security rule is responsible, ask them to explain how they can allow the legitimate crawler without disabling the protection for other automated traffic.
For additional checks on technical issues that can arise during website setup, see the pre-launch website checklist.
6. Check whether the 403 comes from Cloudflare or your origin server
One of the most useful troubleshooting steps is to identify which system generated the 403 response.
A request to a website may pass through multiple layers:
Where a 403 error can occur
Trace the request to find which layer is denying access.
If Cloudflare is returning the 403, review its security events and rules. If the origin server is returning the error, inspect the hosting configuration and logs instead. This distinction can save considerable time because changing the wrong security layer will not resolve the problem.
Once you have identified the source of the error, check for other technical issues that may affect access to your pages. The pre-launch website checklist covers important technical checks that can help prevent avoidable problems.
If you have server access and Cloudflare is proxying your website, you can compare the response received through Cloudflare with a carefully controlled request directly to the origin server. Only do this when you know the correct origin IP and can safely access it; do not expose the origin server or bypass security controls unnecessarily.
7. Test your website's response to AI crawlers
Once you have checked the security settings, test the affected URL from your own computer.
On Windows, you can use curl.exe to avoid the PowerShell alias for curl:
curl.exe -I https://example.com/
This checks the response headers for the requested URL. Replace https://example.com/ with your actual website address.
To test with an AI crawler's user-agent string, you can use:
curl.exe -I `
-A "Mozilla/5.0 (compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot)" `
https://example.com/
For a different crawler, replace the user-agent string with the appropriate documented value.
| HTTP status | What it generally means | Next step |
|---|---|---|
| 200 OK | The server returned the requested resource successfully. | Confirm that the response contains the expected page content. |
| 301 / 302 | The URL redirects to another location. | Check the redirect destination and whether the final URL is accessible. |
| 403 Forbidden | A server or security layer denied the request. | Inspect the relevant security events and server logs. |
| 404 Not Found | The requested resource was not found. | Check the URL, routing, and whether the page still exists. |
| 429 Too Many Requests | The request was rate-limited. | Review request frequency and the configured rate limits. |
| 500 / 503 | A server-side error or service availability issue occurred. | Inspect server health, error logs, and service status. |
You can also use a crawler to inspect your site's URLs and identify other technical problems. See how to find and fix broken links with a website audit for another example of investigating website errors systematically.
Remember that changing the user-agent string does not make your test request equivalent to a real AI crawler request. The request still comes from your own IP address and may be treated differently by security systems that verify IP ranges, signatures, or other traffic characteristics.
For a more complete test, compare your request with the actual crawler's failed request in your server or Cloudflare logs. For OpenAI crawlers, use the official published IP ranges and crawler documentation when checking whether traffic is authentic.
8. A step-by-step troubleshooting checklist
Use the following checklist when an AI crawler receives a 403 error on a website that should be publicly accessible.
- Identify the affected crawler and the URLs returning 403.
- Check
robots.txtfor accidental disallow rules. - Review Cloudflare Security Events for blocked requests.
- Check AI Crawl Control and any Block AI Bots setting.
- Inspect WAF, bot protection, rate limits, and custom firewall rules.
- Determine whether Cloudflare or the origin server generated the 403.
- Review hosting logs, ModSecurity events, and
.htaccessrules. - Ask the hosting provider to investigate if server-level controls are inaccessible.
- Make a narrowly scoped security adjustment for verified, legitimate crawlers.
- Retest the affected URLs and monitor new security events.
For a broader review of your website, use the SEO, GEO, and AEO audit checklist to assess other technical and content-related factors that affect search and AI discovery.
9. Common mistakes to avoid
Resolving a 403 error should not come at the expense of your website's security. These are some mistakes to avoid during troubleshooting.
A crawler can be permitted by robots.txt but still be denied by a firewall, server rule, or application security check. Review both crawl directives and actual HTTP responses.
User-agent strings can be forged. Prefer documented IP ranges, verified-bot features, or other appropriate authentication mechanisms where supported.
Turning off bot protection may resolve a false positive but can also expose your website to unwanted traffic. Identify the specific blocking rule and make a targeted adjustment.
Broadly changing file permissions, deleting .htaccess rules, or disabling ModSecurity can introduce security vulnerabilities. Investigate the relevant logs and consult your hosting provider before making risky changes.
A normal browser and an AI crawler may trigger different security rules. Test the affected crawler's access separately and check the actual request logs.
10. How to prevent AI crawler 403 errors in the future
Once you have resolved the immediate issue, establish a process for monitoring your website's accessibility to legitimate crawlers.
- Review security rules regularly. New firewall rules, bot protection updates, and hosting changes can introduce false positives. Revisit your configuration after significant security or infrastructure changes.
- Monitor server logs. Look for repeated 403 responses from known, legitimate crawler traffic and investigate unexpected patterns. Keep useful timestamps and request identifiers for troubleshooting.
- Keep crawler policies explicit. Document which AI crawlers you permit and which you intentionally block, including the difference between search and training crawlers.
- Coordinate CDN and hosting settings. If you use both Cloudflare and shared hosting, keep track of security rules at both layers. A change in one place may not resolve a restriction in another.
- Retest after infrastructure changes. Changes to hosting, DNS, CDN configuration, or firewall settings can affect crawler access. Test key public URLs after major changes.
- Use the correct response codes. Avoid returning 403 for ordinary rate limiting. Use 429 when a client is being asked to slow down, as recommended by Google for its crawler.
For a complete technical review, the technical SEO audit checklist can help you assess other site-level issues alongside crawler access.
Final thoughts
A 403 error affecting AI crawlers is often a security configuration issue rather than a problem with the content itself. Cloudflare, shared hosting, server firewalls, and application-level protections can all prevent a crawler from reaching a publicly available page.
The first step is to determine where the error originates. Check the crawler's user agent and request logs, review Cloudflare's security events if you use it, and inspect your hosting provider's logs when the origin server is responsible.
After identifying the rule causing the denial, make a targeted change that allows the legitimate crawler while preserving protection against unwanted automated traffic. Then retest the affected URLs and continue monitoring access.
Regularly checking your website's technical health can help you catch access problems before they affect discovery. Explore SiteAuditLint's features to learn more about technical SEO auditing, or visit the SiteAuditLint blog for more practical information on website crawling, auditing, and search visibility.
References
For more details on crawler policies, Cloudflare settings, and HTTP error handling, consult the following official documentation: