GEO & AI search
What Is GEO (Generative Engine Optimization) and How Is It Different from SEO?
A plain-language explanation of generative engine optimization: where the term comes from, how AI search systems retrieve and cite pages, what Google says about it, which crawlers to allow, what you can and cannot influence, and how to measure it.
Key takeaways
- GEO is the work of making your content findable, readable and quotable for AI systems that answer questions by searching the web and summarizing what they find.
- Google says there are no extra requirements for AI Overviews or AI Mode: optimizing for its generative AI features is still SEO. The differences lie in other engines' crawlers and in measurement.
- AI search expands one question into several sub-queries (query fan-out), so a page competes on the sub-questions a buyer has, not only on the original query.
- You can control crawler access, content quality and the consistency of your facts. You cannot control which sources a model picks for a given user, and nobody can guarantee a citation.
- Measure with Search Console's generative AI reports, ChatGPT referral tags in analytics, crawler logs and a fixed set of test prompts run on a schedule.
Generative engine optimization (GEO) is the work of making your content findable, readable and quotable for AI systems that answer questions by searching the web and summarizing what they find: ChatGPT search, Perplexity, Google AI Overviews and AI Mode, Claude and others. The term comes from a 2023 research paper. For most companies GEO is SEO plus a few extra checks: access for non-Google crawlers, pages that answer the sub-questions buyers ask, facts that are easy to quote and verify, and new ways to measure visibility. Google itself says that optimizing for its AI features is still SEO; the real differences lie with the other engines and in measurement.
What is generative engine optimization (GEO)?
GEO means improving the chance that an AI search system retrieves your page and uses it, with a link or a mention, in the answer it writes. Classic SEO aims for a ranked blue link. GEO aims for a place inside the answer. The two overlap heavily, because most AI search systems first run ordinary web searches and then write from the results.
The term was coined in the paper GEO: Generative Engine Optimization by Pranjal Aggarwal and colleagues at Princeton University and IIT Delhi, first published in November 2023 and presented at KDD 2024. The authors built a benchmark of 10,000 queries and tested how rewriting a source changes its visibility in generated answers.
Three findings are useful for businesses:
- Evidence helped most. Adding quotations, statistics or citations to credible sources raised visibility by roughly 30% to 40% on one of the paper’s two visibility metrics. On Perplexity, the live engine they also tested, quotations gave a 22% improvement and statistics up to 37%.
- Keyword stuffing did not work. Adding more search terms gave little or no improvement.
- Smaller sites gained more. Citing sources raised visibility by 115% for pages ranked fifth in the underlying search results, far more than for pages already ranked first.
Read these as directional. The experiments used a test engine built on GPT-3.5 and one live engine in 2023, and today’s systems work differently. Other names for the same field include AEO (answer engine optimization), LLMO and “AI SEO”.
How do AI search systems find and cite sources?
AI search systems retrieve before they write. They turn your question into one or more searches, fetch pages from an index, pick the passages that answer each part, and write a response with links to the pages they relied on. Google calls the search step “query fan-out”: the system issues several related searches across subtopics and combines what they return.
Google described fan-out when it launched AI Mode in May 2025: the system splits a question into subtopics and runs many searches at the same time. Its 2026 guide gives an example. A question about fixing a lawn full of weeds fans out into searches on the best lawn herbicides, removing weeds without chemicals and preventing weeds. A page can be cited for one of those sub-queries even if it would never rank for the original question.
How the main systems get their sources:
- Google AI Overviews and AI Mode use Google’s own index and ranking systems. A page must be indexed and eligible to show with a snippet. Google’s AI features documentation states that there are no additional requirements and no special optimizations needed.
- ChatGPT search uses OpenAI’s crawler OAI-SearchBot. OpenAI’s crawler overview says sites that opt out of OAI-SearchBot are not shown in ChatGPT search answers.
- Perplexity uses PerplexityBot to surface and link websites in its answers. Perplexity states that this crawler is not used to train AI models.
- Claude uses Claude-SearchBot for search and Claude-User when a person asks it to open a page.
How is GEO different from SEO?
From Google’s point of view, not much: its guide to generative AI features says optimizing for them is optimizing for Search, and thus still SEO. The differences appear elsewhere: more crawlers to manage, answers built from passages rather than whole pages, visibility without a click, and measurement that standard rank tracking does not cover.
| Classic SEO | GEO | |
|---|---|---|
| Goal | A ranked link that people click | Being retrieved and cited or mentioned inside an AI answer |
| What competes | Whole pages for a typed query | Passages, facts and brands for the sub-queries behind a question |
| Queries | The words the user typed | Fan-out sub-queries, often more specific, sometimes in another language |
| Crawlers to allow | Googlebot, Bingbot | Also OAI-SearchBot, PerplexityBot, Claude-SearchBot and user-triggered fetchers |
| Content that wins | Relevant, useful, well linked pages | The same, plus direct answers, original data, clear sources and dates |
| Off-site signals | Links, reviews, local listings | Mentions in sources the engines trust, consistent company facts across the web |
| Measurement | Rankings, clicks and impressions in Search Console | Search Console generative AI reports, AI referral traffic, prompt sampling, crawler logs |
| Visibility without a click | Featured snippets | Common: the answer may satisfy the user in place |
The table shows a shift in emphasis, not a separate discipline. A site with crawl problems, thin pages or inconsistent facts will struggle in both.
Which AI crawlers should you allow in robots.txt?
Allow the crawlers that power search answers if you want to be cited, and decide separately about crawlers that collect training data. Blocking a training crawler does not remove you from that company’s search. Check that your CDN or firewall does not block these bots either; Google and OpenAI both mention this.
| User agent | Operator | Purpose | If you block it |
|---|---|---|---|
| Googlebot | Search, including AI Overviews and AI Mode | You disappear from Google Search | |
| Google-Extended | Gemini model training and grounding in Gemini Apps and Vertex AI | No effect on Google Search inclusion or ranking | |
| OAI-SearchBot | OpenAI | ChatGPT search | Your site is not shown in ChatGPT search answers |
| GPTBot | OpenAI | Training of OpenAI models | Your content is excluded from training |
| ChatGPT-User | OpenAI | Pages a user asks ChatGPT to open | May have no effect: robots.txt rules may not apply to user actions |
| PerplexityBot | Perplexity | Surfacing and linking sites in Perplexity answers | Less visibility in Perplexity |
| Perplexity-User | Perplexity | Pages fetched for a user’s question | Usually no effect: it generally ignores robots.txt |
| Claude-SearchBot | Anthropic | Search results in Claude | Less visibility in Claude’s search answers |
| ClaudeBot | Anthropic | Training of Anthropic models | Your future content is excluded from training |
| Claude-User | Anthropic | Pages fetched when a user asks Claude | Claude cannot open your pages for its users |
Names and behavior change. The operators’ documentation is listed in the sources below; check it before you edit your robots.txt.
What can you influence, and what not?
You can influence access, content and consistency: whether crawlers can reach your pages, whether each page answers a question clearly with facts a model can quote, and whether your company facts match everywhere. You cannot influence which sources a model picks for a particular user, how it words the answer, or when the systems change.
Within your control:
- Crawl access and indexing, including CDN and firewall rules.
- Answer-first pages built around the questions buyers actually ask.
- Original material: your own data, prices you actually quote, worked examples, test results with a method and a date.
- Identical company facts on your site, in your Google Business Profile, in directories and in structured data.
- Real third-party presence: reviews, listings, articles and partner directories.
Outside your control: personalization, the variation between two runs of the same prompt, model updates and the commercial choices of each platform.
Google’s 2026 guide also lists tactics you can skip for Google Search. It ignores llms.txt and similar AI text files. Splitting content into tiny chunks is unnecessary, there is no need to write in a special style for AI, and chasing inauthentic mentions “isn’t as helpful as it might seem”. Creating separate pages for every fan-out variation to manipulate AI answers breaks Google’s scaled content abuse policy. Other services may read llms.txt, so it is harmless, but it does not replace good pages.
How do you measure GEO?
Combine four sources: Google’s generative AI reports in Search Console, AI referral traffic in your analytics, your server logs for crawler visits, and a fixed set of test prompts that you run on a schedule. None of them is complete on its own, and no third-party tool sees inside Google’s ranking systems, as Google itself points out.
- Search Console. Google launched generative AI performance reports on 3 June 2026 and extended them to all websites on 31 August 2026. They show impressions in AI Overviews, AI Mode and Discover’s AI features, by page, country, device and date.
- Referral traffic. ChatGPT adds
utm_source=chatgpt.comto links it sends, according to OpenAI’s publisher FAQ. Segment AI referrers in your analytics and watch their conversion, not only their volume. - Crawler logs. Count visits by OAI-SearchBot, PerplexityBot and Claude-SearchBot to your key pages. If they never arrive, content work will not help.
- Prompt sampling. Write 20 to 50 questions your buyers ask, run each several times per engine every month, and record whether you are cited or mentioned, which competitors appear and which pages are linked. Answers vary between runs, so look at trends, not single results.
How we run fan-out research
We reverse-engineer the searches that AI engines run behind a question. We start from one buyer question, write five paraphrases, and send each to OpenAI and Google models through their APIs with web search switched on and the location set to Switzerland, in several passes. We log every sub-query, every source retrieved and every source cited. Then we remove single-site lookups, group the queries by sub-topic and turn them into a content plan.
Three patterns from our September 2026 runs on our own topics:
- Location and language shape retrieval. Asked in English how much a business website costs in Switzerland, the GPT model wrote several of its searches in German and limited 8 of its 10 runs to .ch domains. 26 of its 27 citations pointed to Swiss sites, almost all agency price guides.
- Engines ground in different places. For a question about GEO itself, 42 of the GPT model’s 59 searches looked for official documentation from Google, OpenAI and Perplexity, and it also searched for the arXiv paper. Gemini cited marketing blogs instead.
- Retrieved is not cited. Reddit threads were retrieved dozens to over a hundred times per topic and never cited in these runs.
These runs simulate the models through their APIs; they are not Google’s production retrieval. We use them as direction for content, not as rankings.
Where to start
Start with a short audit: crawler access, indexing, the ten questions your buyers ask most, and whether your company facts match across the web. Then fix the pages that should answer those questions, and set up measurement before you change anything else. Our checklist on preparing business content for AI search goes through the content side step by step.
On a communication-API directory we built, every fact carries a source link, an evidence level and the date it was checked. That structure serves readers first and happens to suit AI systems as well. If you want help with the audit or the ongoing work, see our GEO and AI search visibility and SEO services.
Sources
- Aggarwal et al.: GEO: Generative Engine Optimization (arXiv 2311.09735, KDD 2024)
- Google Search Central: AI features and your website (updated 10 December 2025)
- Google Search Central: Optimizing your website for generative AI features on Google Search (updated 10 July 2026)
- Google Search Central Blog: A new resource for optimizing for generative AI in Google Search (15 May 2026)
- Google Search Central Blog: Introducing Search Generative AI performance reports in Search Console (3 June 2026)
- Google: AI in Search, going beyond information to intelligence (AI Mode and query fan-out, 20 May 2025)
- Google Search Central: Google's common crawlers (Google-Extended)
- OpenAI: Overview of OpenAI crawlers
- OpenAI Help Center: Publishers and Developers FAQ
- Perplexity: Perplexity crawlers
- Anthropic: Does Anthropic crawl data from the web, and how can site owners block the crawler?