How does ChatGPT describe your brand?
All articles
GEO fundamentals/

What Is Crawlability? Making Your Site Readable for AI Search

7 minutes read

What Is Crawlability? Making Your Site Readable for AI Search

Before an AI assistant can cite your page, something has to fetch it and read it. If that step fails, nothing you wrote on the page matters. The page does not exist as far as the answer is concerned.

Crawlability is how easily an automated agent can reach a page and read its content. A page is crawlable when a bot is allowed to request it, the server returns it, and the content is present in what comes back. The term comes from SEO, where the agent is Googlebot. In AI search the agents are different, and they are less forgiving.

Crawlability meaning, in plain terms

Three things have to be true for a page to be crawlable:

  1. The bot is allowed in. Your robots.txt, firewall and CDN do not block it.
  2. The page can be found. Something links to it, or it is listed in a sitemap.
  3. The content is in the response. The text is in the HTML the server sends, not loaded later by scripts the bot never runs.

Crawlability is not the same as indexability. Crawlable means a bot can fetch and read the page. Indexable means a search engine is allowed to store it and show it. A page can be crawlable and marked noindex, or indexable in theory and unreachable in practice.

AI assistants reach your content in three ways, and each depends on crawling:

  • Training. Crawlers collect pages that models later learn from.
  • Search indexes. Engines such as ChatGPT search and Perplexity keep their own indexes, or use a partner's, to find pages for an answer.
  • Live fetching. When a user asks a question, the assistant may fetch a page on the spot to read it.

If any of these is blocked, you lose that route into answers. And there is a second problem. Googlebot renders JavaScript. Most AI crawlers generally do not. A page that looks fine in Google can come back as an empty shell to an AI crawler if its content is rendered in the browser.

The AI crawlers to know

Each company runs several bots with different jobs. These are the main ones at the time of writing. Names change, so check each company's documentation before you edit rules.

CompanyBotWhat it is for
OpenAIGPTBotCollecting training data
OpenAIOAI-SearchBotBuilding the ChatGPT search index
OpenAIChatGPT-UserFetching a page when a user asks
PerplexityPerplexityBotBuilding the Perplexity index
PerplexityPerplexity-UserFetching a page when a user asks
AnthropicClaudeBotCollecting training data
AnthropicClaude-SearchBot, Claude-UserSearch and user-requested fetches
GoogleGooglebotSearch, which also feeds AI Overviews and AI Mode
GoogleGoogle-ExtendedA robots.txt control for Gemini training, not a separate crawler

The split matters. You can block training bots and still allow the search and user bots that get you cited. Blocking everything with "AI" in the name removes you from answers as well.

What blocks AI crawlers

robots.txt rules

A broad Disallow: / under User-agent: *, or a copied "block AI bots" list, is the most common cause. Many of those lists include the search bots.

CDN and firewall bot protection

Some CDNs and security tools block or challenge AI crawlers by default. Your robots.txt can say "allow" while the firewall returns a 403 or a challenge page. Check the logs, not only the rules.

Client-side rendering

If your page is a JavaScript app that fetches its content after load, a crawler that does not run scripts sees a title and an empty div. Server-side rendering or static generation fixes this.

Logins and paywalls

Content behind sign-in is invisible. That includes docs, community forums and pricing that require an account.

Content locked in the wrong format

Text inside images, PDFs without a text layer, tabs and accordions that load on click, and infinite scroll with no paginated links all hide content from simple crawlers.

Weak internal linking

A page nothing links to is hard to discover. Orphan pages and pages missing from the sitemap get crawled late or never.

Slow or failing responses

Crawlers give up on timeouts and repeated 5xx errors, and may come back less often.

How to check your crawlability

You do not need a special tool for a first pass.

  1. Read your robots.txt. Open yoursite.com/robots.txt and look for the bots in the table above. Confirm the search and user bots are not disallowed.
  2. Fetch a page as a bot. Request a key page with a crawler's user agent, for example with curl -A "OAI-SearchBot", and check the status code. A 403 or a challenge page means something upstream is blocking it.
  3. Look at the raw HTML. View source, not the rendered page. Search for a sentence from the middle of your article. If it is not in the source, a non-rendering crawler cannot see it.
  4. Check server logs. Filter for the bot names. You will see whether they visit, which pages they request and what status they get.
  5. Check your sitemap. Every page you want cited should be in it, with a current last-modified date.
  6. Ask the engines. Ask ChatGPT or Perplexity a question your page answers precisely. If they cite a competitor's weaker page every time, access may be part of the reason.
  • Allow OAI-SearchBot, ChatGPT-User, PerplexityBot, Perplexity-User, Claude-SearchBot, Claude-User and Googlebot.
  • Decide on training bots (GPTBot, ClaudeBot, Google-Extended) as a separate choice.
  • Confirm your CDN or firewall is not overriding robots.txt.
  • Render primary content on the server.
  • Keep important content out of login walls, images and click-to-load widgets.
  • Link every important page from at least one other page and include it in the sitemap.
  • Return fast 200 responses and fix redirect chains.
  • Use clear headings and schema markup so the content is easy to parse once fetched.

An llms.txt file can point assistants at your most useful pages. Treat it as a helpful extra. It does not replace any item above, and support for it varies by engine.

Crawlable is the entry ticket, not the win

Fixing crawlability gets your pages into the pool of sources. It does not get them chosen. Engines still pick the source that answers the question most directly, and they often name brands based on third-party pages, not on the brand's own site. We cover that side in how to get cited by ChatGPT and Perplexity.

That is also how you confirm a fix worked. Gensiv does not audit your robots.txt or rendering. It measures the result: for the prompts you track, it records which URLs each engine cites, every day. If your pages start appearing in cited sources after you unblock a crawler or move to server rendering, the fix landed. If they still do not, the problem is the content, not the access.

FAQ

What is crawlability in SEO? It is whether search engine bots can reach and read your pages. Without it, a page cannot be indexed or ranked.

What is the difference between crawlability and indexability? Crawlability is about access: can the bot fetch and read the page. Indexability is about permission and quality: may the engine store and show it.

Should I block AI crawlers? That is a business decision. Blocking training bots keeps your content out of model training. Blocking search and user bots keeps you out of AI answers. Decide on each group separately.

Do AI crawlers run JavaScript? Googlebot does. Most other AI crawlers generally fetch the HTML and do not run scripts, so content that only appears after rendering is at risk.

Does blocking GPTBot remove me from ChatGPT answers? Not on its own. GPTBot is for training. ChatGPT search uses OAI-SearchBot, and live fetches use ChatGPT-User. Blocking those two is what removes your pages as sources.

How do I know if AI engines can read my site? Check robots.txt, fetch a page with a bot user agent, read the raw HTML, and look for the bots in your server logs. Then track whether your pages get cited.

Want to see which of your pages AI engines cite today? Get a free AI visibility report.

Share this article

Become the brand AI recommends.