GEO: how to appear and get cited in ChatGPT, Perplexity and Gemini (2026 guide)
Table of contents
The ten blue links are no longer the whole game: ChatGPT, Perplexity, Gemini and Copilot answer by citing pages. And showing up in GEO is not about a magic file; it comes down to three boring, verifiable things: who can crawl you, how citable your content is, and what you measure.
This guide covers each engine's real user-agents, how to decide who you let in, the structure that gets quoted, and the new Search Console report.
What GEO is (and why Google says it is still SEO)
GEO and its cousin AEO (Answer Engine Optimization) are the work of improving your visibility in AI-generated answers. They are not a new channel with secret rules: they are the same discipline applied to one more surface. Google says it plainly in its official guide to optimizing for generative AI search: "From Google Search's perspective, optimizing for generative AI search is optimizing for the search experience, and thus still SEO."
Two concepts explain the mechanics. Grounding through retrieval (RAG): the system retrieves relevant, recently written pages from its index, drafts the answer on top of them and shows the sources; if you are not in the index, there is nothing to cite. And query fan-out: the model fires off related simultaneous queries, so you are not competing for one query but for several sub-queries at once.
How a generative engine differs from a classic search engine
A search engine gives you a position; a generative engine writes an answer and chooses whom to cite, so your page can be the source without being the top result. Google adds a checkable data point in its documentation on AI features and your website: clicks from pages with AI Overviews are higher quality, meaning users are more likely to spend more time on the site. That does not mean more clicks, it means better visits when they happen. And meeting the requirements does not guarantee crawling, indexing or appearing.
Who reads you: each engine's crawlers
This is where 80 percent of the real work lives: every company separates its training bot from its search bot, and mixing them up is the fastest way to drop out of answers by accident.
| Engine | User-agent | What it is used for | Consequence if you block it |
|---|---|---|---|
| OpenAI | GPTBot | Training | Does not control your presence in ChatGPT |
| OpenAI | OAI-SearchBot | ChatGPT search | You do not appear in search answers |
| OpenAI | ChatGPT-User | User requests | May not respect robots.txt; it does not decide whether you appear |
| Anthropic | ClaudeBot | Training | You exclude future material from training |
| Anthropic | Claude-SearchBot | Search index | Anthropic warns it may reduce your visibility |
| Anthropic | Claude-User | User requests | Reduces visibility for user-initiated searches |
| Perplexity | PerplexityBot | Results; no training | Perplexity recommends allowing it to appear |
| Perplexity | Perplexity-User | User fetches | Generally ignores robots.txt: the user asked for it |
Googlebot | Search, including AI Overviews and AI Mode | The one you cannot lose: no index means nothing to cite | |
Google-Extended | Gemini training control and grounding (Vertex AI) | Does not affect inclusion in Google Search or rankings |
The decision in one copy-paste sentence: training is a business decision; search is a visibility decision. Blocking OAI-SearchBot, Claude-SearchBot or PerplexityBot means giving up on being cited; blocking GPTBot or ClaudeBot, or using Google-Extended, protects your content without touching your rankings. Implementation means per-User-agent rules in robots.txt, one per bot: the details are in a correct 2026 robots.txt and, if you start from zero, in what the robots.txt file is.
An expensive, common mistake: blocking at the WAF or CDN. If your firewall stops the bot from reading your own robots.txt, that block does not count as an opt-out. OpenAI and Perplexity publish their IP ranges as JSON, and propagation of changes takes roughly 24 hours.
Check which bots actually reach you
Anyone can claim to be GPTBot, so cross-check user-agent against IP using the official JSON files: the method is in server log analysis for SEO. And suspect the false block: plenty of sites believe they are excluded when their WAF returns 403 to anything with the word bot in the user-agent.
The content that gets cited
Google's number one priority applies to every engine: non-commodity content with a point of view. The guide contrasts "7 tips for first-time homebuyers", which anyone can write, with "why we waived the inspection & saved money", which only someone who lived it can tell. It also adds that you do not need to chop your content into pieces or rewrite it for AI: these systems understand a page covering several topics.
Blocks that stand on their own out of context work best: a direct one or two sentence answer under the heading, explicit figures and dates, the source cited inside the text, and tables for comparable data. Google also mentions quality images and video, structured data that matches the visible text and brand entity consistency. The limit is clear: creating pages for every keyword variation to manipulate AI violates the scaled content abuse policy.
Generic example: before and after
Before: an h2 that says "More information" and six lines mixing three ideas. After: an h2 that asks "How long does delivery take in 2026?", a one-sentence answer, a table of lead times and the last-reviewed date.
Measuring: the Generative AI report and AI referrals
Since August 31, 2026 the generative AI performance report is live for all websites worldwide. It covers impressions in AI Overviews and AI Mode, with dimensions of pages, countries, dates and devices, and limits that matter: impressions only, with no clicks and no queries; data is attributed to the canonical URL; the most recent data is preliminary.
To see it, your property must be included in Search's generative AI features, and that setting is your opt-out control. The clicks that do happen from pages with AI Overviews show up in the overall report, under the "Web" type, as explained in the guide to Google Search Console. Referral traffic is yours to measure: group chatgpt.com, perplexity.ai, gemini.google.com and copilot.microsoft.com into one channel in your analytics (Google Analytics applied to SEO) and test your ten key queries in ChatGPT and Perplexity by hand, with dates.
How to read the report without fooling yourself
Set a baseline and compare equivalent windows: more AI impressions is not more traffic, and a citation is not a click. If the line falls, diagnose before claiming causation, as explained in how to diagnose a traffic drop. And remember Google's warning: no third-party tool has access to internal metrics, so distrust "secret AI rankings".
What you can ignore and the costly myths
- llms.txt: Google states it does not use it and that it neither helps nor harms in Search. Honest nuance: other systems do use it. Optional infrastructure, not SEO.
- Chunking your content: unnecessary. Write for people.
- Rewriting everything for AI: unnecessary; these systems understand synonyms and intent.
- Chasing mentions: unaudited mentions are not the lever they are sold as.
- "Special AI schema": it does not exist; structured data is still useful for rich snippets, nothing more.
- The hack to rank first in ChatGPT: no company documents a citation factor beyond allowing crawling and publishing useful content.
- AI traffic replacing organic traffic: with no official data behind it, treat it as an additional channel worth measuring separately.
Checklist: 10 actions for GEO
- Review your logs and confirm which AI user-agents visit you.
- Decide in writing what you allow: search yes; training depending on your business.
- Implement per-
User-agentrules in robots.txt, one per bot. - Check your WAF: do not block search bots or stop them reading your robots.txt.
- Confirm in Search Console that your site is included in the AI features.
- Set a baseline: Generative AI report impressions plus AI referrals.
- One h2 per idea and a direct one or two sentence answer under each heading.
- Add verifiable dates, figures and sources.
- Publish non-commodity content: first-hand experience and concrete cases.
- Review at 30, 60 and 90 days: AI impressions, referrals and a manual test.
What works in GEO is what already worked in SEO: let the right crawlers reach you, publish something only you can tell, and measure. To strengthen the context these engines use when citing you, start with the guides to topical authority and entity SEO.