Links That Couldn't Be Crawled: Fixing Incorrect URL Formats

AI OVERVIEW

Crawlers read the href, not the click. Typos like https// or www. without a scheme resolve as relative paths and create junk URLs on your site.

A link can look perfectly fine on the page, take you where you expect when you click it, and still show up in an audit as a link that couldn't be crawled.

When the cause is an incorrect URL format, the destination page is usually fine. The problem is the text inside the href attribute. Crawlers don't click buttons or guess what you meant. They read the href, resolve it into a full URL, and request it.

Read hrefExact text in the attribute
→
ResolveAgainst the page URL or <base>
→
RequestFetch the full URL
→
DiscoverPage joins the crawl

If the value is malformed, the crawler either can't build a URL at all or builds the wrong one.

Which Links Can Google Actually Follow?

Google's documentation is specific here. It follows <a> elements with an href containing a resolvable URL.

MarkupFollowed?
<a href="https://example.com/wallets/">Yes
<a href="/wallets/">Yes
<a href="./wallets/">Yes
<a routerLink="wallets">No
<span href="/wallets/">No
<a>No
<a>No

Google also ignores everything after a # when discovering pages, so href="#" and href="#reviews" don't point to anything new. That list explains most warnings: either the href is missing, or it holds something that isn't a usable URL.

The Sneaky One: Typos That Become Internal 404s

When an href doesn't start with a valid scheme like https:, the browser treats it as a relative path and attaches it to the current page's URL.

CURRENT PAGE https://example.com/blog/guide/ HREF WITH A TYPO (missing colon) https//partner.com/offer WHAT THE CRAWLER ACTUALLY REQUESTS example.com/blog/guide/https//partner.com/offer → 404
One missing character turns an external link into a junk URL on your own domain.

Leaving the scheme off entirely does the same thing: href="www.partner.com/offer" resolves to https://example.com/blog/guide/www.partner.com/offer.

Depending on your server, you'll see a 404, a soft 404, or in the worst case a page that returns 200 and creates a crawlable junk URL. If your audit shows strange internal URLs with another domain name in the path, this is almost always the cause.

Some typos go the other way. htps:// is read as an unknown scheme, so there's no relative resolution, just a link nothing can request.

Common Malformed Formats and Their Fixes

What's in the hrefWhat happensFix
https//site.com/pageResolves as a relative path on your domainhttps://site.com/page
www.site.com/pageResolves as a relative path on your domainhttps://www.site.com/page
htps://site.com/pageUnknown scheme, can't be requestedhttps://site.com/page
https://No host, can't be requestedAdd the full URL or remove the link
javascript:void(0)Not a URL, Google won't follow itReal URL, or a <button>
#Points to the current pageReal URL, or a <button>
"" (empty)Resolves to the current pageAdd the intended URL
audit on /blog/seo/Becomes /blog/seo/audit/audit if you meant the root
/products/red shoesUsually encoded by the browser, but fragile/products/red-shoes

A few notes on the less obvious rows:

  • Placeholders# and javascript:void(0) usually come from a component built as a link that behaves like a button, such as a dropdown toggle or modal opener. If it doesn't navigate, make it a <button>. If it does, give it a real URL.
  • Relative paths without a leading slashhref="audit" is valid, but depends on where the page lives. It works on / and breaks on /blog/seo/. Root-relative URLs behave the same everywhere, which makes them the safer default for internal links.
  • The <base> tagIf a page has <base href="..."> in its head, every relative link resolves against that value instead of the page's own URL. A wrong base tag can break every relative link on a template at once.
  • Spaces and special charactersBrowsers usually percent-encode a space as %20, so the link may work. But URLs with spaces are fragile in sitemaps, email, and analytics. Fix the slug rather than relying on encoding.
  • Odd query strings/search?q or /search?=seo isn't malformed. It's valid syntax. Pointless parameters still create extra crawlable URLs, so treat it as URL hygiene, not a crawl error.
  • Links that aren't meant to be crawledmailto: and tel: links aren't web pages. Some tools list them as uncrawlable. That's expected, not something to fix.

How to Fix Them Properly

  • Get the source page, not just the bad URLFor every flagged link you need two things: the page the link is on and the exact href value. The bad URL alone doesn't tell you where to fix it.
  • Read the actual HTMLInspect the anchor, and check the rendered HTML if JavaScript adds links. Hovering shows the resolved URL, not what's written in the href, which is exactly what hides these problems.
  • Decide if it's a page problem or a template problemOne typo in one post: edit the post. The same bad link on 300 pages: it lives in a shared component like the header, footer, navigation, related-posts widget, or product template. Fix it once.
  • Check the destination after fixing the formatA correctly formatted URL can still return a 404, 410, 403, 5xx, or a redirect chain. Fixing the format and fixing the destination are separate jobs.
  • Re-crawlConfirm the malformed links are gone, the corrected URLs are discovered, and the junk internal URLs no longer appear.

Watch URLs built from variables, too:

const url = protocol + domain + path;

If protocol is "https" instead of "https://", every URL from that line comes out as httpsexample.com/page, which then resolves as a relative path. One missing character can create thousands of broken links.

Why This Matters for SEO

Cost 1

Lost discovery

Crawlers find most pages through internal links. A malformed link removes a route to the destination, along with the internal link value it carries.

Cost 2

Junk URLs

Relative-path typos generate 404 URLs on your own domain that crawlers keep requesting.

Neither is dramatic on one page. Spread across a template, both add up.

Finding Malformed Links Across a Site

The SiteAuditLint broken link checker lists links that lead to 4xx and 5xx responses, along with the pages they're on. Malformed links that resolve to junk internal URLs show up as broken links with strange-looking paths.

To find the typo itself, use custom source search on the crawled HTML for patterns like these:

href="https//href="htps://href="www.javascript:void

You'll get every page containing the pattern, which usually points straight at the responsible template.

[Add screenshot here: a custom source search result for one of these patterns from a real crawl.]

After the fix ships, compare the new audit against the previous one to confirm the broken links are marked as fixed rather than simply moved somewhere else.

Checklist

  • Every navigational link is an <a> element with an href.
  • External URLs start with https:// (colon and both slashes).
  • Internal links use root-relative paths unless there's a reason not to.
  • No # or javascript: values on elements that should navigate.
  • No <base> tag pointing somewhere unexpected.
  • Slugs contain no spaces.
  • Shared templates and URL-building code have been checked.
  • Fixed links return a 200 at their destination.
  • The site has been re-crawled to confirm the fix.
Start with the href. Most of these problems are a missing colon, a missing slash, or a link that was never meant to be a link.