See the agency plan →

Crawlability.ai
Crawlability

What Is Crawlability?

Crawlability Research, AI Search Research, Crawlability.ai··Updated ·5 min read

Crawlability is whether an automated reader — a search engine or an AI engine — can reach a page on your site and read what's on it.

That's the whole definition. A page is crawlable if a bot can request the URL, get served the content, and parse it. If any link in that chain breaks, the page is uncrawlable, and nothing downstream works: it won't be indexed, it won't rank, and it won't be cited in an AI answer.

It's a binary at the page level and a spectrum at the site level. Most sites have some pages that are perfectly crawlable and others that are quietly invisible.

What it isn't

Crawlability gets used loosely, so it's worth separating from the things it sits next to.

It isn't indexability. A crawler can read a page and the engine can still decline to index it. Crawling is access; indexing is a decision made afterward.

It isn't ranking. A page can be crawled, indexed, and still appear on page eight.

It isn't content quality. A crawler doesn't judge. It either gets the bytes or it doesn't.

Crawlability is the first gate. Clearing it guarantees nothing. Failing it guarantees everything downstream fails too.

What actually blocks a crawler

Most crawlability problems are dull and fixable. In rough order of how often they're the cause:

robots.txt disallows the path. The most common and the easiest to miss, because a line written years ago for a staging directory can quietly cover something live.

The server doesn't return the page. A 404 on a page that should exist, a 500 under crawler load, a redirect chain that loops, or a redirect to something unrelated.

The content is drawn by JavaScript after the page loads. The crawler requests the URL, receives an almost-empty HTML shell, and that's what it reads. Some crawlers execute JavaScript, many don't, and the ones that do may not wait long enough. If the text isn't in the HTML source, assume it isn't being read.

It's behind a login, a paywall, or a consent wall. If a human has to click or sign in before the content appears, a crawler sees whatever was there before the click.

Nothing links to it. A page with no internal links pointing at it and no sitemap entry is reachable only if someone already knows the URL. Crawlers follow links; an orphan page waits forever.

The site blocks it on purpose, by accident. Rate limiting, bot protection and firewall rules catch legitimate crawlers alongside bad traffic, and nobody notices because nothing visibly breaks.

Why the word matters again

Crawlability is an old idea. It mattered in 2005 for the same reason it matters now: something automated has to be able to read your page before anything good can happen.

What changed is that there are now two sets of readers with different rules.

Before

Search engine crawling

Googlebot and its equivalents. Well documented, fairly tolerant of JavaScript, and you can see exactly what they did in Search Console. Decades of accumulated convention about how to work with them.

After

AI engine crawling

A separate set of crawlers run by the AI engines, with their own user agents, their own robots.txt tokens, and far less tolerance for pages that assemble themselves in the browser. No console, no crawl report, no notification when something breaks.

A site can be perfectly crawlable by Google and substantially invisible to the AI engines. The two most common reasons: the site renders its content client-side, which Google has largely learned to handle and the AI crawlers often haven't; and the robots.txt blocks AI user agents that nobody on the team has heard of.

Both are invisible from inside. Traffic doesn't drop. Rankings don't move. Nothing appears in any report, because no report exists.

How to check it

In rough order of effort:

View source. Already covered above, and it's the highest-value thirty seconds you can spend.

Read your robots.txt properly. Not just whether it exists — what each group actually allows, and whether any of the AI crawler user agents are disallowed intentionally or by a wildcard somebody wrote years ago.

Check your server logs for crawler hits. If a given crawler has never requested a page, that page is not in its index. Logs are the only honest record of what was actually fetched.

Fetch a page as a crawler would. Request the URL with no JavaScript execution and see what comes back. That's what most AI crawlers get.

The part people skip

Crawlability is checked once, during a site build, and then never again.

Then someone adds a consent banner. A framework upgrade changes rendering from server-side to client-side. A security team tightens the firewall. A robots.txt gets a new line. Each change is reasonable on its own and none of them are announced as crawlability changes, because nobody thinks of them that way.

The symptom is always the same: nothing breaks, nothing errors, and the site quietly stops being readable by something that used to read it.

That's why it's worth re-checking on a schedule rather than when something looks wrong. By the time it looks wrong, it's been wrong for months.

See how AI engines see your brand.

Measure how ChatGPT, Gemini, Claude and Perplexity find, understand and cite you.

7-day trial with 100 credits · card required · cancel anytime