LLM SEO: how large language models choose their sources, and what you can influence

LLM SEO is the work of making your pages and your brand easy for large language models to find, trust and cite. Assistants pick sources in two ways: from what the model absorbed in training, and from live web searches they run while answering. You can't edit the first quickly, but you can influence the second. This guide explains how that selection works on each platform and what to do about it.
- Assistants answer from two places: memory from training data, and pages retrieved live from a search index.
- Retrieval is where you have leverage. If you aren't indexed and crawlable, you can't be cited.
- Platforms disagree: in an Ahrefs study, 76% of AI Overview citations came from Google's top 10, but only about 8% of ChatGPT's did.
- Assistants rewrite one question into several searches, so pages that answer the follow-up questions get more chances.
- Off-site mentions correlate more strongly with AI visibility than backlinks do, so your brand's footprint matters as much as your pages.
What is LLM SEO?
LLM SEO is the work of making your content and your brand easy for large language models to find, trust and cite. You'll also see it called LLM optimization, generative engine optimization (GEO) or answer engine optimization. The labels differ. The goal is the same: when a buyer asks ChatGPT, Gemini, Perplexity or Google's AI Mode about your category, your product shows up in the answer and is described correctly.
Most guides on this topic jump straight to a list of tactics. That's backwards. If you don't know how an assistant decides which pages to read, you can't tell which tactics matter and which are noise. So this guide starts with the mechanics, then turns them into a plan.
If you want the measurement side first, our guide to measuring LLM visibility covers prompts, metrics and reporting. This post is about the inputs: what shapes the answer before anyone measures it.
The two places an answer comes from
Every answer an assistant gives draws on two kinds of knowledge.
- Model memory (training data). During training, the model reads a huge slice of the public web and forms associations: which brands belong to which category, what they're known for, who they compete with. This is why ChatGPT can name project management tools without searching. You can't change this memory quickly. It updates only when a new model is trained, and you have no direct say in what it absorbs.
- Live retrieval (search). For many questions, the assistant runs web searches while answering, reads a handful of pages, and writes the answer from them. OpenAI says ChatGPT "may search the web automatically when your question would benefit from current information." Google says its AI Overviews and AI Mode ground their answers in pages retrieved by its core Search systems.
Retrieval is where you have leverage. A page you publish or fix this week can be retrieved next week. It's also where most LLM SEO advice quietly assumes the action is, without saying so.
The two paths also feed each other. Pages and mentions that exist today are candidates for tomorrow's training data. So work that improves retrieval also improves memory over time, just more slowly.
How retrieval picks sources, step by step
The platforms don't publish their ranking code, but Google and OpenAI describe enough of the process to sketch it. Here's the sequence, with what each vendor has said about it.

- Decide whether to search. ChatGPT searches when it judges that current information would help. Questions that need current or specific facts, such as vendors, pricing and comparisons, are more likely to trigger a search than broad definitions.
- Rewrite and fan out the question. The assistant turns one prompt into several search queries. OpenAI's help page shows a single question becoming "one or more targeted queries" sent to search partners. Google calls this query fan-out: it generates "concurrent, related queries" to fetch more results for the same question.
- Retrieve from a search index. Google says its AI features use "our core Search ranking systems to retrieve relevant, up-to-date web pages." ChatGPT uses its own crawler, OAI-SearchBot, plus third-party search providers.
- Select passages. From the retrieved pages, the model picks the specific passages that answer the question. Pages that state a clear answer, with the entity, the claim and the evidence close together, are easier to use.
- Write and cite. The model writes the answer and links some of the pages it used. Not every page that shaped the answer gets a citation, and not every brand mentioned gets a link.
Why ChatGPT, Perplexity and Google cite different pages
If retrieval runs on search, you'd expect AI citations to mirror Google rankings. For Google's own AI features, they largely do. For standalone assistants, they mostly don't.
Ahrefs checked 15,000 long-tail queries in July 2025 and measured how many AI-cited URLs also ranked in the top 10 for the same query. The results varied by platform:

| Platform | Where it retrieves from | Overlap with Google top 10 | What it means for you |
|---|---|---|---|
| Google AI Overviews / AI Mode | Google's index and ranking systems, with fan-out | 76% | Classic SEO is the main lever. Rank for the fan-out queries too. |
| Perplexity | Live web search, built to cite | 28.6% | Strong rankings help, but it reaches beyond page one. |
| Gemini | Google search, triggered selectively | 8.6% | Cites less often from page one. Brand footprint matters. |
| Copilot | Bing | 8.2% | Check Bing Webmaster Tools, not only Search Console. |
| ChatGPT | OAI-SearchBot plus third-party search providers | 8% | Allow OAI-SearchBot. Your Google rank alone won't carry you. |
Two cautions on those numbers. First, they come from one study at one point in time. Ahrefs' larger March 2026 analysis of 863,000 keyword SERPs found only 37.9% of AI Overview URLs also appeared in the top 10, so the AI Overview figure moves with the sample and method. Second, overlap isn't causation. The direction is still useful: Google's AI features stay close to Google rankings, and ChatGPT goes its own way.
For platform-specific detail, see our guides on how ChatGPT picks SaaS brands and ranking in AI Overviews and AI Mode.
What actually influences selection
Put the mechanics together and five factors stand out. Some are documented by the platforms; others come from correlation studies, which show association, not proof.
1. Access and indexing (documented)
Google says a page must be "indexed and eligible to be shown in Google Search with a snippet" to appear in its AI features. OpenAI says sites that block OAI-SearchBot won't appear in ChatGPT's search answers, though it may still show a bare link and title it found elsewhere. Note that OpenAI's crawlers are separate: blocking GPTBot (training) doesn't block OAI-SearchBot (search). Many SaaS sites block the wrong one.
2. Relevance to the rewritten queries
Because of fan-out, your page competes on the sub-questions, not only the prompt a buyer typed. Pages that cover pricing, integrations, use cases, alternatives and limits for a clear audience give the assistant more to retrieve.
3. Extractable answers
Models lift passages, not whole pages. A paragraph that names your product, says what it does for whom, and backs the claim with a number or source is easy to reuse. The GEO research paper presented at KDD 2024 found that content changes could raise visibility in generative answers by up to 40%, with results that varied by domain. Adding citations, quotations and statistics were among the methods it tested.
4. What the rest of the web says about you
Ahrefs studied 75,000 brands and found branded web mentions had the strongest correlation with AI Overview visibility (0.664). Backlinks came in at 0.218. For model memory, this is even more important: the model learned your category from third-party pages, not from your homepage. See how to earn AI brand mentions for the practical side.
5. Freshness
Retrieval favors pages that look current for time-sensitive questions. Vercel's team, who wrote about adapting their SEO for LLMs, refresh content on 30, 90 and 180-day cycles. They also reported that ChatGPT referred around 10% of new Vercel signups at the time of writing, in June 2025.
What doesn't matter as much as you'd think: special files and markup. Google's July 2026 guidance says you don't need llms.txt, AI-specific schema or content "chunking" to appear in its AI features. Our llms.txt guide covers what the file does and doesn't do.
LLM SEO vs traditional SEO
LLM SEO is not a replacement for SEO. It sits on top of it. Crawlability, indexing and helpful pages are still the foundation, especially for Google's AI features. What changes is the target and the scoreboard.
| Traditional SEO | LLM SEO | |
|---|---|---|
| Goal | Rank a page and earn the click | Be read, cited and described accurately |
| Unit of competition | A keyword | A prompt and its fan-out queries |
| Main off-site signal | Backlinks | Brand mentions and third-party descriptions |
| Crawlers to allow | Googlebot, Bingbot | Also OAI-SearchBot and other assistant bots |
| How you measure | Rankings, clicks, conversions | Mention rate, citation rate, accuracy across a prompt set |
The search results for "llm seo" reflect the same shift. When we pulled the US SERP in Ahrefs on 29 September 2026, the page opened with an AI Overview citing a Reddit thread and a Vercel blog post, followed by Reddit discussions and a YouTube course. Community content and brand-owned case notes sat above most "ultimate guides." Search demand is real, too: Ahrefs shows about 1,600 US searches a month for "llm seo" and 1,700 for "llm optimization," both with a keyword difficulty of 1.
See your own AI visibility, free
Get a one-page audit of how ChatGPT, Perplexity and Gemini treat your brand against up to three competitors, plus your biggest gap. In your inbox within 48 hours.
Get my free auditA 6-step LLM SEO plan for B2B SaaS
Here's the order we'd work in. Each step maps to a part of the selection process above.
- Fix access (week 1). Check robots.txt for OAI-SearchBot, Googlebot and Bingbot. Confirm your pricing, product and comparison pages are indexed in Search Console and Bing Webmaster Tools. Make sure key content renders without JavaScript.
- Build a prompt set (week 1). Write 20 to 40 prompts your buyers actually ask, from "what is" to "best X for Y" to "X vs Y." Run them in ChatGPT, Perplexity and Google AI Mode. Record who gets mentioned, who gets cited and what's said about you. Our AI visibility tracking guide has a template.
- Map the fan-out (week 2). For each priority prompt, list the sub-questions an assistant would need to answer: pricing, integrations, security, alternatives, fit by company size. Check which ones your site answers clearly.
- Write extractable pages (weeks 2 to 6). Fill the gaps with pages that state the answer early, name the audience, and support claims with numbers or sources. Refresh dated pages rather than publishing thin new ones.
- Grow your footprint (ongoing). Get listed and described correctly on review sites, comparison articles, partner pages and relevant communities. Correct outdated descriptions where you can.
- Measure monthly. Rerun the same prompt set. Answers vary from run to run, so look at trends across the set, not single responses.
Common LLM SEO mistakes
- Blocking the search bot while trying to block training. Blocking GPTBot is a valid choice. Blocking OAI-SearchBot removes you from ChatGPT's search answers.
- Chasing one prompt. Assistants give different answers to the same prompt. One screenshot is an anecdote, not a baseline.
- Writing for machines. Stuffed keywords and robotic FAQ blocks don't help. Google says plainly that you don't need to rewrite content for AI systems.
- Ignoring off-site descriptions. If review sites describe your product as it was three years ago, assistants will too.
- Treating it as separate from SEO. For Google's AI features, rankings still do most of the work. Don't pull budget from SEO basics to fund "GEO hacks."
FAQ
What is LLM SEO?
LLM SEO is the practice of making your content and brand easy for large language models like ChatGPT, Gemini and Perplexity to find, understand and cite. It overlaps heavily with classic SEO, because most assistants retrieve pages from a search index before answering. The difference is that you also care about how your brand is described across the web, not only where your pages rank.
How is LLM SEO different from traditional SEO?
Traditional SEO aims for a ranking position and a click. LLM SEO aims to be one of the few sources an assistant reads and cites, and to be described accurately when it doesn't cite you. Crawl access, indexing and helpful content still matter. Third-party mentions, clear positioning and answers to follow-up questions matter more than they used to.
How do I improve LLM SEO?
Start by confirming crawlers like OAI-SearchBot and Googlebot can reach your key pages and that those pages are indexed. Then build pages that answer the specific questions buyers ask, including the follow-ups, and earn mentions on the review sites, comparisons and communities assistants read. Measure a fixed set of buyer prompts monthly so you can see what changes.
Do I need special markup or an llms.txt file for LLM SEO?
Not for Google. Its July 2026 guidance says you don't need llms.txt, special schema or rewritten content to appear in AI Overviews or AI Mode. Structured data is still useful for regular search features, and llms.txt is harmless, but neither is the main lever.
Sources
- Google Search Central: Optimizing your website for generative AI features on Google Search (updated July 2026)
- OpenAI Help Center: ChatGPT search
- OpenAI Help Center: Publishers and developers FAQ
- OpenAI: Overview of OpenAI crawlers
- Ahrefs: Only 12% of AI cited URLs rank in Google's top 10 (Linehan and Guan, August 2025)
- Ahrefs: AI Overview citations and the top 10 (Linehan, March 2026)
- Ahrefs: brand signals and AI Overview visibility, 75,000 brands (May 2025)
- Aggarwal et al.: GEO: Generative Engine Optimization (KDD 2024)
- Vercel: How we're adapting SEO for LLMs and AI search (June 2025)
- Ahrefs Keywords Explorer and SERP overview, US, pulled 29 September 2026


