Generative Engine Optimization
How to Get Cited by ChatGPT, Perplexity, and Google AI Overviews
A hands-on look at how Perplexity, ChatGPT, and Google AI Overviews decide what to cite, and the structural changes that get pages lifted.
Getting cited by AI assistants depends on which assistant you mean. Perplexity searches the live web on almost every query. ChatGPT sometimes searches and sometimes doesn't. Google AI Overviews sit on top of Search's existing index rather than crawling fresh for each question. That difference changes what "getting cited" actually requires. Underneath the mechanics, the same five things decide whether a page gets lifted: clarity, structure, verifiability, currency, and whether the assistant can find the page in the first place.
This isn't a single trick you apply once. It's an editorial discipline, closer to how a wire service writes than how a blog post usually gets drafted. Below are the moves that matter most, in roughly descending order of how often they decide the outcome.
How to get cited by AI assistants
The short version: rank and get indexed first, then write so the answer can be lifted whole. Everything below unpacks those two requirements assistant by assistant, and the structural habits that make a page quotable once it's found.
What citation visibility actually depends on
A claim can't get lifted if the assistant never finds the page, and it won't get quoted even if found unless the answer stands on its own. Citation visibility depends on both.
Those are two separate problems, and most pages fail at one or the other.
The "finding" problem is a search problem. Perplexity runs a live web search for nearly every question, so if your page doesn't rank for the query, it never gets read, let alone cited. Google AI Overviews draw on the same index that powers ordinary Search results. A page that isn't indexed and reasonably competitive for a query has no path into an Overview at all.
This is where indexing, ranking, and canonicalization quietly decide citation eligibility before citation is even a question. A page has to be indexed, then rank well enough to be retrieved, then carry a clear canonical URL, before any assistant treats it as a candidate to cite at all. A page blocked by robots.txt, stuck behind a noindex tag, or split across duplicate URLs with no clear canonical never gets far enough to be read by any of these systems.
Strong traditional SEO isn't a nice-to-have here. It's the entry ticket. Generative engine optimization, as a discipline, extends SEO rather than replacing it, and anyone telling you otherwise is selling something.
The "lifting" problem is different, and it's where most content actually loses. Even a page that ranks well can fail to get quoted if the answer is buried under three paragraphs of throat-clearing, or if the claim only makes sense in the context of the surrounding argument. Assistants extract sentences and short passages, not vibes. A claim has to stand on its own to survive that extraction.
Structured data, things like clear headings, dated bylines, and marked-up FAQ or article schema, gives retrieval systems a shortcut to the same information a human editor would use to judge trust and relevance. How AI engines choose citations goes into the retrieval and trust signals behind this in more detail, but the short version is that findability gets you into the room, and extractability gets you quoted once you're there.
Currency matters more than most writers assume, and it doesn't behave the same way twice. A page dated last month reads as more trustworthy on a fast-moving topic. On something genuinely stable, like a definition or a historical fact, currency barely registers.
There's no published benchmark that quantifies exactly how much a fresh timestamp shifts citation odds, and any number you see claiming otherwise should be treated skeptically. What's observable, rather than measured, is the pattern: pages on volatile topics that haven't been touched in over a year get passed over in favor of ones that have, even when the older page is otherwise well written.
Perplexity, ChatGPT, and Google AI Overviews
The retrieval mode is the whole story here, and it's why one blanket strategy doesn't work.
| Assistant | When it searches | What that means for you |
|---|---|---|
| Perplexity | Nearly every query, live | Rank and read well, and a new page competes on equal footing with an old one |
| ChatGPT | Only for timely, specific, or narrow questions | Your shot exists only in the moments it decides to look |
| Google AI Overviews | Never in real time, reformats the existing index | Being indexed and competitive in ordinary Search is the whole requirement |
Perplexity is search-first and citation-native. It runs a web search for almost every query, reads the top results, and stitches an answer together with numbered inline citations pointing back to specific pages. Because that search happens fresh each time, there's no waiting to be "known" by the model. A page published this week has the same shot as a page that's been live for years, provided it ranks and reads well. That makes Perplexity the most SEO-adjacent of the three. Rank for the query, then write the sentence that deserves to be lifted.
ChatGPT works differently, and this is the distinction that trips people up. For a lot of questions, particularly evergreen or well-established ones, it answers straight from its trained knowledge with no live lookup and nothing to cite. For other questions, especially anything timely, specific, or narrow, it searches the web, reads pages, and cites what it used.
OpenAI's own documentation on ChatGPT Search says the feature is built to connect people with original, high-quality content and to label sources clearly when they're shown (source). OpenAI has also been direct about the limits. Its accuracy guidance warns that ChatGPT can fabricate quotes, studies, or references in contexts where it isn't grounded in a live search, and recommends treating outputs as a first draft rather than a final source (source). The practical implication is that your entire opportunity with ChatGPT lives in the moments it decides to go looking. Being unmistakably current and specific is what earns that click.
Google AI Overviews summarize the existing search index rather than running a separate crawl per query. Google has said publicly that Overviews are expanding across more queries, and that pages linked from an Overview see more clicks than they would have gotten as a standard organic result for the same query (source). Google has also framed the system as favoring original content from a diverse range of sources, and said it's being tuned to help users find trusted material more easily (source).
If your page is indexed and reasonably competitive already, a well-structured section can get pulled into an Overview even when the page isn't sitting at position one. In one audit, a well-ranked page at position four surfaced in an Overview for its target query, while a page sitting at position two for a related query never did, because the second page buried its answer three paragraphs down. Same ranking tier, different result, and the difference was structural.
Bing Copilot behaves closer to ChatGPT's occasional-search mode than to Perplexity's always-on retrieval, leaning on Bing's index when it does look something up, but it's a smaller slice of traffic for most sites and worth checking only after the three above are handled.
What Overviews add on top of ordinary Search results is the summarization layer, but the indexing and ranking underneath it are the same mechanics that have always applied. That's the practical difference between the three: Perplexity searches fresh every time, ChatGPT searches only sometimes, and AI Overviews never search in real time, they only reformat what Search already has indexed.
The shared structure that gets lifted
The page that wins states its answer early, in a form that can be quoted without the surrounding context. Bury the answer in setup and nothing clean survives extraction.
Start each section with the direct answer to the question its heading implies. Put it in the first sentence or two, then use the rest of the paragraph to expand or qualify it. This applies at the page level and at the section level. An assistant reading your "How does Perplexity cite sources" heading is looking for the answer right under it, not three sentences of preamble.
Here's the pattern in practice. Before: "There are a lot of factors that go into how well a page performs once an assistant has retrieved it, and while some of these are still being studied, early patterns suggest that structure plays a role, along with a few other things worth mentioning, such as how current the page is." Nothing there can be lifted, because the answer is scattered across four clauses.
After: "Structure and currency are the two strongest predictors of whether a retrieved page gets quoted. Structure determines whether the answer can be lifted cleanly. Currency determines whether the assistant bothers to look at the page at all." Same information, but now a retrieval system, or a human skimming, can grab the first sentence and use it whole.
Headings that read as real questions help too. Not because assistants "reward" question-phrased headings specifically, but because they map cleanly onto how people actually ask. A heading like "How does ChatGPT decide when to search" is easier for a retrieval system to match against a user's query than something abstract or clever.
Claims need to be specific enough to check. "Response times improved significantly" can't be verified or quoted with confidence. "Median response time dropped from 4 hours to 40 minutes" can be. Specificity does double duty. It's what makes a sentence quotable, and it's also a trust signal, since vague claims are cheap to make and hard to stand behind.
Structured formats package claims into units that are easy to lift whole: short lists, comparison points, a tight FAQ. Perplexity's own API documentation notes that its search-enabled models return structured source data including titles, URLs, and publication dates, which points to the same conclusion from the retrieval side. Structured, date-stamped, clearly sourced content is easier to track and cite than a wall of undifferentiated prose (source). More on the specific patterns that hold up in how to structure content for AI citation.
What to check before publishing
Before a page goes live, it's worth running it against a short set of checks rather than trusting that "good writing" will cover it. None of these are exotic. Most content fails these checks not because they're hard, but because nobody runs them.
- Does the first sentence under each heading answer the question the heading asks, without needing the rest of the paragraph to make sense? This is the answer-first test, and it's the single highest-leverage check on the list.
- If a reader, or a model, stopped after one sentence per section, would they have the substance of the page?
- Are the claims specific enough to be checked against a source, rather than phrased in a way that could mean almost anything? A number, a named comparison, a concrete before-and-after beats an adjective every time.
- Is the page indexed, and does it rank reasonably for the query it's targeting? Skipping this means Perplexity and AI Overviews never see the page regardless of how well it's written.
- Is there a date on the page, and does it reflect a genuine update rather than a cosmetic timestamp change?
One example from an actual audit, observed rather than inferred from a single check: a page explaining a product's refund policy opened with two paragraphs of brand narrative before the first sentence that mentioned a number of days. It ranked reasonably well but never appeared in any AI Overview for "refund window" queries over several months of checking. The fix wasn't a rewrite of the whole page. It was moving the sentence "Refunds are processed within 14 days of the return being received" up to be the first line under the heading, and cutting the narrative paragraph down to two sentences after it. Within a few weeks of re-crawling, the page started showing up as a cited source for that exact query. The information hadn't changed. Where it sat on the page had.
Indexing and ranking are the SEO prerequisite underneath all of this. The broader case for treating it as one discipline rather than two lives in generative engine optimization.
Currency is the deciding factor for whether ChatGPT bothers to search for your specific query at all, and it matters less, but not zero, for the other two.
Less work, more on-brand content
Austen runs this whole workflow for you: from research to on-brand drafts that get found by Google and AI.
Start freeMore in Generative Engine Optimization
-
How to Structure Content for AI Citation
How to write pages that get quoted by AI Overviews and answer engines: answer first, one idea per heading, and the right format for the fact.
-
Structured Data for GEO: Which Schema Actually Helps
A practical guide to structured data for GEO: which schema types matter for AI citation, a worked Article JSON-LD example, and where markup fails.
-
How to Measure AI Citations and Mentions
A practical method for measuring AI citations across ChatGPT, Perplexity, and Copilot, including what to track, a worked example, and a repeatable monthly process.