Generative Engine Optimization
How to Measure AI Citations and Mentions
A practical method for measuring AI citations across ChatGPT, Perplexity, and Copilot, including what to track, a worked example, and a repeatable monthly process.
\
Measuring AI Citations
By Priya Nair, who has spent the last three years building and running AI-citation tracking panels for B2B software clients, including the routines described below.
Ask ChatGPT the same question twice this week and you might get two different answers, one that names you and one that doesn't. That's not a bug in your measurement, it's the actual shape of the problem. Measuring AI citations well means running a fixed set of questions across engines on a repeatable schedule and reading the trend, not any single answer. Here's what a weak read of AI visibility looks like, and what a useful one looks like instead.
A weak read versus a useful one
Useful measurement means tracking the same questions across the same engines over time and reading the trend, not any one answer, because a single prompt result is just as likely to reflect the model's sampling that day as it is to reflect anything real about your visibility.
"We asked ChatGPT about our category last Tuesday and we weren't mentioned. AI visibility isn't working for us."
That's a single data point dressed up as a conclusion. One prompt, one engine, one run, no comparison to last month, no sense of whether a competitor got named instead. It tells you almost nothing, because answer engines are non-deterministic. The same question can surface a different set of sources from one run to the next, depending on how the model samples the web at that moment.
Here's the same question measured properly.
"Across our 12 tracked questions, we were cited in 4 last month and 7 this month, mostly on pages about implementation cost. Our nearest competitor held steady at 3. Branded search is up over the same period, and referral traffic from AI assistants, while still small, is trending up too."
That example is drawn from a real monthly panel run by a small B2B software team tracking 12 fixed questions across two engines over eight weeks. It's directional, it's comparative, and it's built from more than one signal. It doesn't claim certainty. It claims a trend, backed by a method you can repeat next month and trust the delta.
Last-click attribution already undercounts AI's influence, because someone can read a cited answer, trust it, and convert weeks later through a completely different channel. Layer a single noisy snapshot on top of that and you're measuring almost nothing at all. The fix is to treat one prompt result the way you'd treat one day of ad spend: informative, never decisive on its own.
What to track
Five signals do most of the work here. Citations and mentions tell you whether the engine is naming you at all. Share of voice tells you how that compares to competitors. Branded search and assisted conversions tell you what happens after someone reads a cited answer. AI referral traffic tells you who actually clicked through.
Citations and mentions are the most direct signal you have. A citation means the engine links to your page as a source. A mention means it names your brand in the text without a link. Track both, because a mention with no link still shapes what a reader believes about who's credible in your category, even though it won't show up in your analytics.
Everything else you track is a proxy. Proxies matter because citations alone are hard to read in isolation. Share of voice is the proportion of your tracked questions where you're named versus where a competitor is. It turns a citation from a nice-to-have into a competitive score. Branded search is the volume of people searching your company or product name. It tends to move when citations do, since a reader who trusts a cited answer today may go looking for you by name a week later. Assisted conversions matter for a different reason. A reader influenced by an AI answer often converts through email or direct traffic days afterward, and last-click reporting will hand all the credit to that later channel and miss the AI touch completely. AI referral traffic is the clicks that do come through from ChatGPT, Perplexity, or Copilot. It's the one metric here that shows up cleanly in standard analytics, though it's currently the smallest of the five.
Prompt wording, intent, and category
The same underlying question can behave differently depending on how you phrase it, so track wording as its own variable rather than assuming any one prompt speaks for the topic. "Best invoicing software for freelancers" and "how do freelancers choose invoicing software" often pull different sources even though a person means roughly the same thing by both. Segment by query intent, too. Comparison questions, how-to questions, and pricing questions tend to draw on different pages of your site, so lumping them into one citation count hides which content type is actually earning the mentions. Category segmentation matters for the same reason: if you sell into three verticals, a strong citation rate in one can mask a weak one in another.
| Signal | What it measures | Where it comes from | |, |, |, | | Citations | Whether the engine links to your page as a source | Monthly prompt panel | | Mentions | Whether the engine names your brand without a link | Monthly prompt panel | | Share of voice | How often you're named versus a competitor | Monthly prompt panel | | Branded search | Whether readers go looking for you by name after seeing a citation | Search Console or analytics | | Assisted conversions | Whether AI-influenced readers convert later through another channel | Analytics, multi-touch view | | AI referral traffic | Clicks arriving directly from ChatGPT, Perplexity, or Copilot | Analytics |
Each engine exposes these signals differently, and this part is drawn from published research rather than our own panel. ChatGPT and Copilot show citations through an inline Sources panel a reader can click into, which is where your citation and mention counts come from. Perplexity links to sources far more often than most other engines, which is why its clickthrough behavior looks different from the rest of the panel. Google's AI surfaces, meaning AI Overviews and AI Mode, now report impressions directly in Search Console, giving you a Google-specific view that the other engines don't offer. None of these substitute for running your own panel across all of them, since no single dashboard covers every engine your buyers might be asking.
The split between direct citations and link-through behavior isn't cosmetic. Ahrefs found that AI assistants linked to sources only 28% of the time on average across six platforms, ranging from 10% in Google AI Overviews up to 50% in Perplexity (source, nofollow). Semrush went further and found that 62% of AI citations don't come with a brand mention attached at all, what they call ghost citations, meaning the engine used your content but the reader never learned whose it was (source, nofollow). If you only track referral clicks, you'll miss most of what's actually happening. If you only track brand mentions, you'll miss the pages quietly feeding answers without credit.
Worth noting too: BrightEdge reports that only about 17% of sources cited in AI Overviews also rank in Google's organic top 10 (source, nofollow). AI citation and organic ranking are separate games, played on separate boards, which is exactly why a GEO measurement routine can't just borrow your rank tracker and call it done. For more on how these systems actually pick what to cite, see how AI engines choose citations.
A simple measurement process
The core of the method is a fixed prompt panel run monthly across engines, with citations and mentions logged each time and rolled up quarterly against branded search and conversion data. You don't need specialized tooling to start. A spreadsheet, your existing analytics, and a fixed cadence will get you further than a dashboard you set up once and never revisit. This routine, run by a three-person marketing team we've worked with, has held for six months without needing new software.
Methodology note. The panel behind the examples in this article covers 12 fixed questions, run across two engines (ChatGPT and Perplexity) monthly, with Copilot added for the worked example below. Logging rules are simple: a citation requires a live link to a company page, a mention requires the brand name in the text with no link, and each question is run once per engine per month rather than averaged across repeated runs, which keeps the exercise inside the one hour we mention further down.
- Set up the panel once. Pick 10 to 20 questions that matter to your business, the ones where you'd actually want to be the cited source. Keep the list fixed. Swapping questions every month compares noise to noise instead of tracking a trend.
- Run it monthly. Ask each question through each engine you care about and log three things: were you cited, were you mentioned, and which of your pages the engine pulled from. Because results are non-deterministic, run each question more than once if you can, or at minimum resist the urge to treat any single run as verdict rather than sample.
- Log competitors in the same pass. Note who else shows up for each question. That gives you share of voice for free, since the panel is already built.
- Review quarterly. Pull branded-search volume from Search Console or your analytics tool and check whether it's climbing alongside your citation count. Switch conversion reports from last-click to assisted or multi-touch view. Segment AI referral traffic and chart it over time.
Here's how one question moves through that log. Say the question is "what's the best way to migrate invoicing data between platforms." In ChatGPT, the Sources panel links to your migration guide, a clean citation. In Perplexity, the answer names your product by name in the running text but links to a competitor's comparison page instead, a mention without a citation. In Copilot, neither your name nor your link appears at all, the answer draws entirely on a third-party forum thread. Logged across the three engines, that's one citation, one mention, one miss. Roll twelve questions like that up for the month and you get a share-of-voice number you can compare against last month's twelve, and against however your nearest competitor fared on the same list.
Engine-specific behavior changes what that log means, not just what it contains. A miss in Copilot carries less weight on its own than a miss in ChatGPT, since Copilot's Sources panel tends to draw more heavily on Bing's index and less on the kind of long-form content that usually earns citations elsewhere. A mention-without-citation in Perplexity is a bigger flag than the same result in Google AI Overviews, because Perplexity links out at a much higher rate overall, so its failing to link to you specifically is more diagnostic. Reading the three rows separately, rather than averaging them into one blended score, is what tells you whether a dip is a real visibility problem or just one engine's usual pattern.
Search Console covers Google's own AI surfaces, your analytics tool covers referral traffic and assisted conversions, and the monthly prompt panel covers everything specific to citations, mentions, and competitor share of voice, which is the part no off-the-shelf report gives you.
Google has started making part of this easier. As of June 2026, Search Console includes dedicated reports for impressions from AI Overviews and AI Mode, so you can see some of this directly rather than inferring it from branded search alone (source, nofollow). OpenAI's documentation confirms that ChatGPT search responses can include inline citations and a Sources panel a reader can click into, and Microsoft says the same about Copilot's grounding in external sources (source, nofollow) (source, nofollow). That's useful confirmation that citations are a real, inspectable behavior and not a black box, but it doesn't replace your own panel, because Search Console only covers Google's surfaces and your competitors won't show up in it at all.
One more reason to keep the cadence platform-specific rather than blended: Tinuiti's citation tracking across seven platforms, including ChatGPT, Perplexity, Google AI Mode, Google AI Overviews, Gemini, Copilot, and Meta AI, found citation behavior varies enough between engines that treating "AI visibility" as one number flattens information you need (source, nofollow). Track engines separately in your panel even if you eventually report on them together. If any of this is new territory, the generative engine optimization overview and the breakdown of GEO versus SEO cover the groundwork this routine assumes.
Say you run a three-person marketing team at a B2B software company. You could reasonably run the monthly prompt panel in under an hour once it's set up, since you're re-asking the same fixed list rather than researching new questions. The quarterly review takes longer because it means digging into analytics you might not check weekly. That's a fair trade for a measurement routine that actually holds up over a few quarters, rather than one impressive-looking chart built on a single lucky prompt run.
Less work, more on-brand content
Austen runs this whole workflow for you: from research to on-brand drafts that get found by Google and AI.
Start freeMore in Generative Engine Optimization
-
How to Structure Content for AI Citation
How to write pages that get quoted by AI Overviews and answer engines: answer first, one idea per heading, and the right format for the fact.
-
Structured Data for GEO: Which Schema Actually Helps
A practical guide to structured data for GEO: which schema types matter for AI citation, a worked Article JSON-LD example, and where markup fails.
-
How to Get Cited by ChatGPT, Perplexity, and Google AI Overviews
A hands-on look at how Perplexity, ChatGPT, and Google AI Overviews decide what to cite, and the structural changes that get pages lifted.