AI Search

LLM SEO: how to get cited by large language models

By Nihanth Guntur · 2026-09-29

LLM SEO is the work of getting large language models such as ChatGPT, Claude and Perplexity to find, read and cite your pages. It comes down to two jobs: let the AI search crawlers each vendor documents actually reach your site, then write passages clear enough that a model can lift them as an answer. In our audits, the first job is the one that fails without anyone noticing.

What is LLM SEO, and how is it different from ordinary SEO?

LLM SEO, sometimes called LLM optimization, SEO for LLMs or LLMO, is ordinary SEO plus one extra question: can each AI system's own crawler fetch your pages, and does the page contain a passage worth quoting? The ranking fundamentals do not go away.

Google says this plainly. Its guide to generative AI features in Search states that “optimizing for generative AI search is optimizing for the search experience, and thus still SEO”, and that AI Overviews and AI Mode rely on its core ranking systems to retrieve pages from the Search index. So for Google, LLM SEO is SEO: a page must be indexed and eligible to show with a snippet.

ChatGPT, Claude and Perplexity are different in one practical way. Each runs its own search crawler, with its own user agent and its own robots.txt token. If that crawler cannot fetch your page, the page cannot be retrieved for an answer on that platform, however well it ranks on Google. That is why we frame LLM SEO as two jobs rather than one.

Job one: which crawlers does LLM SEO depend on?

LLM SEO depends on the search crawlers, not the training crawlers. Each vendor documents them separately, and each lets you allow search while opting out of training.

PlatformSearch crawler (allow for citations)Training crawler (separate choice)User-triggered fetcher
ChatGPT (OpenAI)OAI-SearchBotGPTBotChatGPT-User
Claude (Anthropic)Claude-SearchBotClaudeBotClaude-User
PerplexityPerplexityBotNone listed; its docs say PerplexityBot is not used for foundation modelsPerplexity-User
Google AI Overviews and AI ModeUses the Search indexNot covered by the guide we linkNot covered by the guide we link

OpenAI's crawler overview says OAI-SearchBot “is used to surface websites in search results in ChatGPT’s search features”, and that sites opted out of it “will not be shown in ChatGPT search answers, though can still appear as navigational links.” It also says each setting is independent: you can allow OAI-SearchBot and disallow GPTBot. OpenAI adds that it can take about 24 hours after a robots.txt change for its search systems to adjust.

Anthropic's help centre article says Claude-SearchBot analyses content to improve the relevance and accuracy of search responses, and that disabling it “may reduce your site's visibility and accuracy in user search results.” Restricting ClaudeBot, by contrast, signals that future materials should be excluded from training datasets.

Perplexity's crawler documentation says PerplexityBot is designed to surface and link websites in its search results and is not used to crawl content for AI foundation models. Perplexity recommends allowing it in robots.txt and permitting requests from its published IP ranges.

Why is crawler access the part of LLM SEO that fails silently?

Because robots.txt is only one gate. A firewall, bot-management rule or CDN setting can refuse the request before robots.txt is ever consulted, and a tool that only reads robots.txt can report the site as open.

What we have seen

On one SEO Autopilot engagement, a US healthcare analytics company on WordPress behind a firewall, the firewall returned 403 to OpenAI's and Anthropic's search crawlers on every page. The trigger was a generic bad-bot rule matching the substring “searchbot”. Eleven other crawlers passed. The client's monitoring tool reported all sixteen AI crawlers as accessible, because it read robots rules instead of sending requests.

Between the baseline frozen on 15 July 2026 and the re-measurement on 18-19 August 2026, across all 39 sitemap URLs, the AI crawler access score went from 45 to 82 and overall AI Visibility from 18.35 to 66.35. There was a cost: Performance fell from 48 to 28, because removing the firewall rule took the CDN edge cache with it (median time-to-first-byte 0.42s to 2.39s). Crawler access is a configuration change with trade-offs, not a checkbox.

The vendors point at the same risk. Perplexity's documentation says that if you use a web application firewall, you may need to explicitly whitelist its bots, and recommends combining user-agent matching with IP verification. Anthropic's article notes that blocking its IP addresses may not work correctly as an opt-out, because it impedes their ability to read your robots.txt.

Send real requests, do not just read the rules. Here is the procedure we use:

  1. Read your robots.txt and confirm there is no Disallow for OAI-SearchBot, Claude-SearchBot or PerplexityBot, either by name or through a wildcard group.
  2. Decide separately on GPTBot and ClaudeBot. Blocking training crawlers is a legitimate choice. OpenAI's overview says its settings are independent, and Anthropic documents ClaudeBot and Claude-SearchBot as separate bots.
  3. Request a sample of pages, not only the homepage, with each search crawler's user agent string, and record the HTTP status. A 403 or 429 is a block, whatever robots.txt says.
  4. Search your firewall and bot-management rules for substrings such as “bot”, “crawler” or “searchbot” that could match these user agents by accident.
  5. Check your server logs over the following days for requests from the real crawlers, and whether they received a 200.

For a quick first look, our free AI visibility checker reads robots.txt rules for 11 AI crawlers and requests your homepage once with each crawler's user agent. That is the checker's own request, not the real crawler's, and it covers the homepage only, so it cannot prove a crawler reaches your site. It can show you a block worth investigating.

Job two: how do you write passages an LLM will cite?

Write every section so that its first sentence or two answer one specific question on their own. A model retrieving your page is looking for a passage it can quote; a passage that depends on three paragraphs of context is hard to lift.

This is our practice, not a vendor rule, and it is the checklist we write to:

  • Open the page with a direct answer to the question in the H1, in plain words, before any background.
  • Phrase several H2s as the questions a reader would actually type, and answer each one in the first line below it.
  • Keep one idea per paragraph, and name the subject in each passage rather than relying on “it” or “this” from earlier text.
  • Use tables for comparisons and numbered steps for procedures, so the structure survives extraction.
  • State only figures you can source, and link the source in the same sentence.
  • Give every page exactly one H1 that says what the page is about.

That last point is less obvious than it sounds. On the same healthcare analytics site, 21 of 39 pages had no H1 because the post template rendered the title as an H2. One template change took pages with exactly one H1 from 9 of 39 to 30.

Do not overbuild on markup. Google's guide says structured data is not required for generative AI search and there is no special schema.org markup to add, though it still helps with rich results. It also says Google Search ignores llms.txt files, while noting that creating them is fine for other services that use them.

What does LLM SEO not require?

LLM SEO does not require a new site, a separate content set, or tricks aimed at models. It requires crawlable pages and clear answers.

In practice that means you should not rewrite pages for machines, stuff keywords into answer paragraphs, or buy a tool that promises AI mentions. No one can promise that a model will cite you. What you can control is whether each platform's search crawler gets a 200, and whether the page, once fetched, contains a passage that answers the question better than the alternatives. If you want this run as a managed service, that is what our AI SEO services cover, alongside ordinary technical and on-page SEO.

For the full crawler list, including user-triggered fetchers, see our guide to every AI crawler and how to allow it. For the firewall case in more depth, read AI crawlers blocked by a firewall.

Frequently Asked Questions

What is LLM SEO?

LLM SEO is the practice of getting large language models such as ChatGPT, Claude and Perplexity to find and cite your pages. It combines letting their search crawlers reach your site with writing passages clear enough to be quoted as an answer.

Is LLM SEO different from SEO?

For Google, no: its guide says optimising for generative AI search is still SEO. For ChatGPT, Claude and Perplexity, you also need their own search crawlers, OAI-SearchBot, Claude-SearchBot and PerplexityBot, to be able to fetch your pages.

Should I block GPTBot and ClaudeBot?

That is a separate decision from LLM SEO. OpenAI and Anthropic document their training crawlers separately from their search crawlers, so you can block GPTBot and ClaudeBot for training while still allowing OAI-SearchBot and Claude-SearchBot for search.

Do I need an llms.txt file for LLM SEO?

Not for Google: its guide says Google Search ignores llms.txt files, and that creating one neither harms nor helps rankings there. Other services may use such files, so creating one is optional.

How do I know if AI crawlers are blocked on my site?

Request your pages with each search crawler's user agent and check the HTTP status, then look for the real crawlers in your server logs. Reading robots.txt alone can miss a firewall rule that returns 403.

Free AI Visibility Check

See whether AI search crawlers reach your homepage

Our free checker reads robots.txt rules for 11 AI crawlers and requests your homepage once with each crawler's user agent. It is the checker's own request, not the real crawler's.