Crawlability vs Indexability
These two words get used as if they mean the same thing. They don't, and the difference decides what you should actually go and fix.
A page can be read but not stored. It can be stored but never surfaced. Treating all three as one problem is how teams spend a month on the wrong fix.
Three states, not two
Crawlable
A bot can request the URL, gets served the content, and can parse it. This is access. It's binary — either the bytes arrive and are readable, or they don't.
Indexed
The engine kept what it read and can retrieve it later. This is a decision the engine makes after crawling, and it can decline for reasons that have nothing to do with access.
Cited
The engine surfaced your page in an answer and attributed it. This is a competition. Being indexed only makes you eligible to enter it.
Each state depends on the one before it, and none of them guarantees the next. Most diagnosis failures come from assuming they do.
Crawlable but not indexed
This is the most common gap and the most misdiagnosed.
The crawler reached your page, read it, and the engine decided not to keep it. Reasons it does that:
- A
noindexdirective in the page head or the HTTP headers - A canonical tag pointing somewhere else, so the engine treats this URL as a duplicate
- The content is near-identical to another page it already has
- There's very little on the page to index
- The engine has limited patience for the site and spent it elsewhere
None of these are access problems. Opening up robots.txt, adding sitemap entries, or improving internal linking fixes none of them — the crawler was never the obstacle.
The tell: your server logs show the crawler fetching the page, and the page isn't in the index. Access worked. The decision went against you.
Indexed but not cited
The newer gap, and the one with the least established practice around it.
Your page is in the index. Someone asks an engine a question your page answers. The engine names a competitor instead.
Being indexed means you're eligible. It doesn't mean you win. The engine chose something it judged more useful, more current, more clearly structured, or more corroborated elsewhere.
This one is uncomfortable because the fix isn't technical. There's no directive to add. The page has to be a better answer, or the brand has to be more substantiated outside its own website.
Not crawlable at all
The simplest state and the easiest to miss, because nothing about it looks broken from the inside. The page loads fine in a browser. Humans use it. It just isn't being read. The underlying mechanics are covered in what crawlability is.
Telling them apart
You need two sources, and they answer different questions.
Server logs tell you about crawling. Did the bot request this URL, and what did your server return? This is the only honest record of access. If a crawler has never requested a page, nothing about indexing is relevant yet.
The index tells you about indexing. For Google, Search Console's URL inspection answers it directly. For the AI engines there is no equivalent console, which is the whole difficulty — you can only infer indexing by asking questions and seeing whether you get cited.
Nothing tells you about citation except asking. Repeatedly, because the answers vary between runs. One query is an anecdote.
Which to fix first
In order, because each gate is wasted effort if the one before it is shut:
- Access. Can the crawler get the page and read the body text? Check the raw HTML, the robots.txt and the server response. Nothing else matters until this is true.
- Indexing. Is there a
noindex, a canonical pointing elsewhere, or a near-duplicate? These are usually one-line fixes with an immediate effect. - Citation. Only once the first two are clean. This is content and authority work, it's slow, and it's impossible to measure if the page wasn't readable in the first place.
Teams routinely start at three. It's the most interesting problem and the one everyone wants to talk about. But improving a page that no crawler can read is work nobody will ever see.
Why three states, not two
The classic framing is crawl, then index. That was enough when the only question was whether you appeared in a list of links.
Appearing in an AI answer is a different outcome from being in an index, and it can fail independently. A site can be flawlessly crawled, fully indexed, and still absent from every answer in its category — and the old two-state model has no way to describe that, let alone diagnose it.
Three states is the minimum needed to say where the problem actually is.