The Reframe

LLM SEO: how large language models choose their sources, and what you can influence

By , The Reframe··10 min read
Cover: LLM SEO, how AI picks its sources

LLM SEO is the work of making your pages and your brand easy for large language models to find, trust and cite. Assistants pick sources in two ways: from what the model absorbed in training, and from live web searches they run while answering. You can't edit the first quickly, but you can influence the second. This guide explains how that selection works on each platform and what to do about it.

Key takeaways
  • Assistants answer from two places: memory from training data, and pages retrieved live from a search index.
  • Retrieval is where you have leverage. If you aren't indexed and crawlable, you can't be cited.
  • Platforms disagree: in an Ahrefs study, 76% of AI Overview citations came from Google's top 10, but only about 8% of ChatGPT's did.
  • Assistants rewrite one question into several searches, so pages that answer the follow-up questions get more chances.
  • Off-site mentions correlate more strongly with AI visibility than backlinks do, so your brand's footprint matters as much as your pages.

What is LLM SEO?

LLM SEO is the work of making your content and your brand easy for large language models to find, trust and cite. You'll also see it called LLM optimization, generative engine optimization (GEO) or answer engine optimization. The labels differ. The goal is the same: when a buyer asks ChatGPT, Gemini, Perplexity or Google's AI Mode about your category, your product shows up in the answer and is described correctly.

Most guides on this topic jump straight to a list of tactics. That's backwards. If you don't know how an assistant decides which pages to read, you can't tell which tactics matter and which are noise. So this guide starts with the mechanics, then turns them into a plan.

If you want the measurement side first, our guide to measuring LLM visibility covers prompts, metrics and reporting. This post is about the inputs: what shapes the answer before anyone measures it.

The two places an answer comes from

Every answer an assistant gives draws on two kinds of knowledge.

Retrieval is where you have leverage. A page you publish or fix this week can be retrieved next week. It's also where most LLM SEO advice quietly assumes the action is, without saying so.

The two paths also feed each other. Pages and mentions that exist today are candidates for tomorrow's training data. So work that improves retrieval also improves memory over time, just more slowly.

How retrieval picks sources, step by step

The platforms don't publish their ranking code, but Google and OpenAI describe enough of the process to sketch it. Here's the sequence, with what each vendor has said about it.

Five-step diagram of how an AI assistant picks sources: decide to search, rewrite and fan out queries, retrieve from a search index, select passages, then write and cite
How retrieval-based AI answers choose their sources. Based on Google Search Central and OpenAI documentation, September 2026.
  1. Decide whether to search. ChatGPT searches when it judges that current information would help. Questions that need current or specific facts, such as vendors, pricing and comparisons, are more likely to trigger a search than broad definitions.
  2. Rewrite and fan out the question. The assistant turns one prompt into several search queries. OpenAI's help page shows a single question becoming "one or more targeted queries" sent to search partners. Google calls this query fan-out: it generates "concurrent, related queries" to fetch more results for the same question.
  3. Retrieve from a search index. Google says its AI features use "our core Search ranking systems to retrieve relevant, up-to-date web pages." ChatGPT uses its own crawler, OAI-SearchBot, plus third-party search providers.
  4. Select passages. From the retrieved pages, the model picks the specific passages that answer the question. Pages that state a clear answer, with the entity, the claim and the evidence close together, are easier to use.
  5. Write and cite. The model writes the answer and links some of the pages it used. Not every page that shaped the answer gets a citation, and not every brand mentioned gets a link.
Tip Step 2 is the one most teams ignore. A buyer asking "best SOC 2 automation tool for a 50-person startup" may trigger searches for pricing, integrations and alternatives. A page that answers only the head question misses every one of those fan-out searches.

Why ChatGPT, Perplexity and Google cite different pages

If retrieval runs on search, you'd expect AI citations to mirror Google rankings. For Google's own AI features, they largely do. For standalone assistants, they mostly don't.

Ahrefs checked 15,000 long-tail queries in July 2025 and measured how many AI-cited URLs also ranked in the top 10 for the same query. The results varied by platform:

Bar chart of the share of AI-cited URLs that also rank in Google's top 10: AI Overviews 76%, Perplexity 28.6%, Gemini 8.6%, Copilot 8.2%, ChatGPT 8%
Share of cited URLs that also rank in Google's top 10 for the same query. Source: Ahrefs, 15,000 long-tail queries, July 2025.
PlatformWhere it retrieves fromOverlap with Google top 10What it means for you
Google AI Overviews / AI ModeGoogle's index and ranking systems, with fan-out76%Classic SEO is the main lever. Rank for the fan-out queries too.
PerplexityLive web search, built to cite28.6%Strong rankings help, but it reaches beyond page one.
GeminiGoogle search, triggered selectively8.6%Cites less often from page one. Brand footprint matters.
CopilotBing8.2%Check Bing Webmaster Tools, not only Search Console.
ChatGPTOAI-SearchBot plus third-party search providers8%Allow OAI-SearchBot. Your Google rank alone won't carry you.

Two cautions on those numbers. First, they come from one study at one point in time. Ahrefs' larger March 2026 analysis of 863,000 keyword SERPs found only 37.9% of AI Overview URLs also appeared in the top 10, so the AI Overview figure moves with the sample and method. Second, overlap isn't causation. The direction is still useful: Google's AI features stay close to Google rankings, and ChatGPT goes its own way.

For platform-specific detail, see our guides on how ChatGPT picks SaaS brands and ranking in AI Overviews and AI Mode.

What actually influences selection

Put the mechanics together and five factors stand out. Some are documented by the platforms; others come from correlation studies, which show association, not proof.

1. Access and indexing (documented)

Google says a page must be "indexed and eligible to be shown in Google Search with a snippet" to appear in its AI features. OpenAI says sites that block OAI-SearchBot won't appear in ChatGPT's search answers, though it may still show a bare link and title it found elsewhere. Note that OpenAI's crawlers are separate: blocking GPTBot (training) doesn't block OAI-SearchBot (search). Many SaaS sites block the wrong one.

2. Relevance to the rewritten queries

Because of fan-out, your page competes on the sub-questions, not only the prompt a buyer typed. Pages that cover pricing, integrations, use cases, alternatives and limits for a clear audience give the assistant more to retrieve.

3. Extractable answers

Models lift passages, not whole pages. A paragraph that names your product, says what it does for whom, and backs the claim with a number or source is easy to reuse. The GEO research paper presented at KDD 2024 found that content changes could raise visibility in generative answers by up to 40%, with results that varied by domain. Adding citations, quotations and statistics were among the methods it tested.

4. What the rest of the web says about you

Ahrefs studied 75,000 brands and found branded web mentions had the strongest correlation with AI Overview visibility (0.664). Backlinks came in at 0.218. For model memory, this is even more important: the model learned your category from third-party pages, not from your homepage. See how to earn AI brand mentions for the practical side.

5. Freshness

Retrieval favors pages that look current for time-sensitive questions. Vercel's team, who wrote about adapting their SEO for LLMs, refresh content on 30, 90 and 180-day cycles. They also reported that ChatGPT referred around 10% of new Vercel signups at the time of writing, in June 2025.

What doesn't matter as much as you'd think: special files and markup. Google's July 2026 guidance says you don't need llms.txt, AI-specific schema or content "chunking" to appear in its AI features. Our llms.txt guide covers what the file does and doesn't do.

LLM SEO vs traditional SEO

LLM SEO is not a replacement for SEO. It sits on top of it. Crawlability, indexing and helpful pages are still the foundation, especially for Google's AI features. What changes is the target and the scoreboard.

Traditional SEOLLM SEO
GoalRank a page and earn the clickBe read, cited and described accurately
Unit of competitionA keywordA prompt and its fan-out queries
Main off-site signalBacklinksBrand mentions and third-party descriptions
Crawlers to allowGooglebot, BingbotAlso OAI-SearchBot and other assistant bots
How you measureRankings, clicks, conversionsMention rate, citation rate, accuracy across a prompt set

The search results for "llm seo" reflect the same shift. When we pulled the US SERP in Ahrefs on 29 September 2026, the page opened with an AI Overview citing a Reddit thread and a Vercel blog post, followed by Reddit discussions and a YouTube course. Community content and brand-owned case notes sat above most "ultimate guides." Search demand is real, too: Ahrefs shows about 1,600 US searches a month for "llm seo" and 1,700 for "llm optimization," both with a keyword difficulty of 1.

See your own AI visibility, free

Get a one-page audit of how ChatGPT, Perplexity and Gemini treat your brand against up to three competitors, plus your biggest gap. In your inbox within 48 hours.

Get my free audit

A 6-step LLM SEO plan for B2B SaaS

Here's the order we'd work in. Each step maps to a part of the selection process above.

  1. Fix access (week 1). Check robots.txt for OAI-SearchBot, Googlebot and Bingbot. Confirm your pricing, product and comparison pages are indexed in Search Console and Bing Webmaster Tools. Make sure key content renders without JavaScript.
  2. Build a prompt set (week 1). Write 20 to 40 prompts your buyers actually ask, from "what is" to "best X for Y" to "X vs Y." Run them in ChatGPT, Perplexity and Google AI Mode. Record who gets mentioned, who gets cited and what's said about you. Our AI visibility tracking guide has a template.
  3. Map the fan-out (week 2). For each priority prompt, list the sub-questions an assistant would need to answer: pricing, integrations, security, alternatives, fit by company size. Check which ones your site answers clearly.
  4. Write extractable pages (weeks 2 to 6). Fill the gaps with pages that state the answer early, name the audience, and support claims with numbers or sources. Refresh dated pages rather than publishing thin new ones.
  5. Grow your footprint (ongoing). Get listed and described correctly on review sites, comparison articles, partner pages and relevant communities. Correct outdated descriptions where you can.
  6. Measure monthly. Rerun the same prompt set. Answers vary from run to run, so look at trends across the set, not single responses.

Common LLM SEO mistakes

FAQ

What is LLM SEO?

LLM SEO is the practice of making your content and brand easy for large language models like ChatGPT, Gemini and Perplexity to find, understand and cite. It overlaps heavily with classic SEO, because most assistants retrieve pages from a search index before answering. The difference is that you also care about how your brand is described across the web, not only where your pages rank.

How is LLM SEO different from traditional SEO?

Traditional SEO aims for a ranking position and a click. LLM SEO aims to be one of the few sources an assistant reads and cites, and to be described accurately when it doesn't cite you. Crawl access, indexing and helpful content still matter. Third-party mentions, clear positioning and answers to follow-up questions matter more than they used to.

How do I improve LLM SEO?

Start by confirming crawlers like OAI-SearchBot and Googlebot can reach your key pages and that those pages are indexed. Then build pages that answer the specific questions buyers ask, including the follow-ups, and earn mentions on the review sites, comparisons and communities assistants read. Measure a fixed set of buyer prompts monthly so you can see what changes.

Do I need special markup or an llms.txt file for LLM SEO?

Not for Google. Its July 2026 guidance says you don't need llms.txt, special schema or rewritten content to appear in AI Overviews or AI Mode. Structured data is still useful for regular search features, and llms.txt is harmless, but neither is the main lever.

Sources

  1. Google Search Central: Optimizing your website for generative AI features on Google Search (updated July 2026)
  2. OpenAI Help Center: ChatGPT search
  3. OpenAI Help Center: Publishers and developers FAQ
  4. OpenAI: Overview of OpenAI crawlers
  5. Ahrefs: Only 12% of AI cited URLs rank in Google's top 10 (Linehan and Guan, August 2025)
  6. Ahrefs: AI Overview citations and the top 10 (Linehan, March 2026)
  7. Ahrefs: brand signals and AI Overview visibility, 75,000 brands (May 2025)
  8. Aggarwal et al.: GEO: Generative Engine Optimization (KDD 2024)
  9. Vercel: How we're adapting SEO for LLMs and AI search (June 2025)
  10. Ahrefs Keywords Explorer and SERP overview, US, pulled 29 September 2026
HumaFounder of The Reframe. An electrical and aerospace engineer turned growth marketer with 10+ years of experience, including work with Fortune 500 tech and SaaS companies.