The advice flooding this space is “get cited, get cited”, as though a citation in an AI answer were a clean, countable outcome you can optimise toward. It is not. Citation depends on the exact query a person types, and conversational queries do not behave like search terms. There is no exact match, no phrase match, no broad match, no stable denominator, so what the tracking tools give you is a sample over a prompt set someone chose, not a measurement of a channel.
This is not an argument to ignore AI search. It is an argument to stop paying for it as a separate, measurable channel when it is mostly good SEO with a new label and a dashboard on top. Google now says the same thing in its own documentation, the major engines barely cite the same sources as each other, and the data keeps pointing back at the same lever: credible third-party coverage, which is PR. Care about AI visibility. Just do not fall for the gimmick layer that sells it as something you can measure cleanly, because you cannot measure it without importing your own bias.
What this article covers
- Why classic search was measurable and conversational AI search is not, in the same way
- The bias built into how citation-tracking prompt sets are designed
- What tools like Profound and Scrunch actually give you, and what they cannot
- Why the useful data kept pointing back at PR and SEO, and what to do with that
Over the last year, at an agency before I founded Ridley Digital, I spent a lot of time on answer engine optimisation. I wrote AEO briefs, ran visibility audits, and built roadmaps to help brands show up in AI answers. I went in willing to believe it was a new discipline. I came out convinced of something quieter and less marketable: that almost everything that works in AEO is good SEO, and that the part the industry has bolted on top, the measurement, the packaging, the sense that this is a separate channel you can track and optimise, is where the story gets oversold.
So this is not a sceptic who never tried it taking a swing from the sidelines. It is the opposite. I did the work, I found real value in some of it, and that is exactly why I want to be honest about the parts that do not hold up. The aim here is not to talk you out of caring about AI search. It is to stop you paying agency rates for a repackaged version of work you should already be doing, on the strength of a number that cannot mean what it is sold to mean.
A year inside the work, and the conclusion it led to
The pattern became hard to ignore. Every AEO engagement I worked on resolved, in the end, into the same recommendations: make the content genuinely useful and specific, structure it so it is easy to extract, get the technical foundations right so it can be crawled and indexed, and earn credible third-party coverage. Strip the labels off an AEO roadmap and an SEO roadmap and you would struggle to tell them apart. The deliverables were SEO deliverables. The audit was an SEO audit. The brief was a content brief.
That is the uncomfortable thing the category does not advertise. A lot of what is sold as answer engine optimisation, or generative engine optimisation, is SEO with a 2.0 sticker on it, priced as a new service because the label is new. There is nothing wrong with the underlying work. The work is good. What is wrong is selling it as a separate discipline with its own budget line, when the thing actually moving the needle is the SEO you were supposed to be doing anyway.
What AEO and GEO actually mean, and where the real difference sits
Before going further it helps to define the terms, because the vocabulary is part of the confusion. Answer engine optimisation, AEO, usually means structuring content so an AI surface can extract a clean, direct answer from it. Generative engine optimisation, GEO, usually means getting a brand cited or recommended inside an AI-generated response. You will also see AIO, LLMO, GSO and others. The proliferation of acronyms is itself a tell: industry analysis has found that a majority of practitioners use the terms inconsistently, and fewer than a third hold to one definition across a year. When a discipline cannot agree on what to call itself, it is usually because the thing underneath is not as new as the labels imply.
There is a real distinction worth keeping, and it is the one that survives scrutiny. Some of this work is on-site: the content, the structure, the technical foundations, the things you control on your own pages. Some of it is off-site: how credibly and consistently other people describe you across the web, which feeds the sources the engines draw on. The on-site half is, almost entirely, SEO by another name. The off-site half is real and matters, but as we will see, it is PR and digital authority rather than a new optimisation science. Holding those two halves apart is the key to seeing what is genuinely new here and what is repackaged.
Why search was measurable and this is not
To see why AI visibility resists measurement, it helps to remember why classic search submitted to it so well. Search has a stable, countable unit: the query. People type a finite, repeatable set of terms, and the whole apparatus of search marketing is built on that. Match types exist because the query is a discrete thing you can pin down. Exact match, phrase match, broad match are all ways of describing how close a real search was to a term you care about. Above all, search has a denominator. Impressions tell you how often you could have been seen, so a click-through rate means something, because it is measured against a real, observed population of searches.
Conversational AI queries have none of that. People do not query an LLM in tidy, repeatable phrases. They ask long, idiosyncratic, context-laden questions, then follow up and refine across the thread of a conversation the model remembers. There is no exact match because there is barely a repeat. There is no phrase or broad match because there is no stable phrase to match against. And there is no denominator: nobody can tell you how many times your brand could have surfaced, because the population of possible conversations is effectively infinite and unobservable. You are not measuring against impressions. You are measuring against a guess.
On top of that, the engines are non-deterministic. Ask the same question twice and you can get two different answers, with two different sets of sources. This is not a complaint about immature tooling; it is in the nature of the systems. Even Profound, one of the better tracking platforms, states plainly that AI platforms do not return the same answer twice. Industry analysts put it the same way: almost every generated response differs from the last, where ten Google searches for the same term give you a reliable read on what Google will show. When the thing you are measuring will not hold still, a single reading of it is a snapshot of one roll of the dice, not a measurement of a stable quantity.
The prompt problem, which is really a bias problem
Here is the part I am least comfortable about from my own time doing this, because it is a flaw I worked inside without fully naming at the time. To track citations, you need a set of prompts to test. The tools let you auto-generate them, upload your own, or pull from a panel of real queries the vendor has collected. On most engagements, we built the prompt set from what we thought users would search for, which is to say we wrote prompts the way we would write target keywords.
That is the bias, and it is baked in before any data exists. We were importing a search-term mindset into a medium that does not behave like search terms. The prompts reflected how we imagined people phrase things, not the messy, conversational, multi-turn way people actually use LLMs, with all the context and follow-up that a flat prompt list cannot capture. So the citation data that came back was not a neutral reading of reality. It was a reading of how well the brand performed against our own assumptions about how people ask. Change the prompt set and the numbers move. When the measurement depends that heavily on the assumptions of the person designing it, you do not have a metric, you have a mirror.
What the tools actually give you
I want to be fair to the platforms, because they are genuinely clever and they do provide something. Tools like Profound and Scrunch monitor a range of answer engines, ChatGPT, Perplexity, Gemini, Copilot, Google’s AI surfaces and more, by running prompts and capturing the responses. They report how often your brand is mentioned, how often a page of yours is cited, how that compares to competitors as a share of voice, the sentiment of how you are described, and which third-party sources the engines lean on. Scrunch and others layer on auditing and content suggestions. As a diagnostic instrument, that is genuinely useful.
But notice what it is and is not. It is a sample of citations across a prompt set you or the vendor chose, scored against competitors on that same set. It is not a measurement of a population, because there is no population to measure. Google itself has now made this point in its own documentation, noting that third-party tools have no access to its internal ranking data, cannot guarantee performance, and that any predictions they offer are their own inferences from the outside rather than readouts from the system. A tool reporting an AI visibility score is modelling, not measuring, and it is worth being clear-eyed about the difference.
Why there is no single number to report
Even setting the prompt bias aside, the headline share-of-voice figure hides a deeper problem: the major engines barely cite the same sources as each other, so a single number averages across systems that fundamentally disagree.
The evidence here is now substantial and consistent across independent studies. A meta-analysis of around 680 million citations found that only about 11 percent of the domains cited by ChatGPT are also cited by Perplexity. That is not a rounding difference, it is two almost entirely separate views of who counts as a source, and a separate study of 118,000 responses arrived at the same 11 percent figure independently. Even within one company the divergence holds: Google’s AI Overviews and its AI Mode have been found to share only around 13.7 percent of cited URLs despite reaching broadly similar answers. And the overlap between AI-cited sources and Google’s own top-ten organic results has reportedly collapsed from roughly 70 percent to under 20 percent in about two years.
Sit with what that means for measurement. If you are dominant in one engine you can be nearly absent in another, and a tool that reads one engine, or blends them into a single score, will never show you that gap. The engines also behave differently in kind, not just in degree: some lean on a brand’s own pages, others pull mostly from third parties, and at least one major engine builds its answers more from knowledge baked into training than from live retrieval, which is why per-page optimisation barely moves it. A single “AI share of voice” number flattens all of that into something that looks like a measurement and behaves like a weather forecast for a climate that changes when you look at it.
The attribution gap: cited, and invisible anyway
There is one more gap that matters if you care about outcomes rather than vanity. You usually cannot close the loop from a citation to a result. Google’s AI Overviews pass no attribution data, ChatGPT only began tagging its outbound links partway through 2025, and a good deal of AI-influenced traffic lands in analytics as direct, with no referrer at all.
The scale of the disconnect shows up in the publishers’ numbers. Major outlets that are cited constantly in AI answers have been found to receive well under one percent of their referral traffic from those AI platforms, because being named in an answer is not the same as being clicked. So even when the citation happens, and even when you can see it in a tracking tool, the line from that citation to a visit, let alone a customer, is mostly invisible. You can be cited and never know it produced anything. That is a strange thing to build a separately-budgeted, ROI-reported channel around.
What the data kept pointing at
For all that, the data was not useless. It was just useful in a way that quietly undercut the premise. The most consistent, actionable thing it told us was that third-party mentions drive visibility. When a brand was cited well, it was usually because credible external sources, coverage, reviews, mentions on respected sites, were saying consistent things about it, and the engines were leaning on those sources. Our best lever for improving AI visibility, again and again, turned out to be earning more and better third-party coverage.
This is not just my impression from a year of engagements. A large analysis of 75,000 brands found that brand web mentions correlate with AI citation rates at around 0.66, roughly three times more strongly than backlinks at around 0.22. In other words, the single biggest measurable driver of whether an engine cites you is how much the rest of the web is talking about you, which is the founding premise of public relations. The AEO data led us, over and over, back to PR.
And PR is already a visibility discipline, with its own decades-old, imperfect but established ways of measuring reach and impact. We were using an expensive new tool to rediscover that being talked about by credible third parties makes you more visible. The insight was real. It just was not new, and it did not need a separate channel called AEO to deliver it. It needed good PR and good SEO, pointed at the sources the engines respect.
Google’s own answer settles a lot of the argument
If this were only my read from one year of engagements, you would be right to weigh it as one practitioner’s opinion. But the largest player in the space has now said the same thing in writing. In its guidance on optimising for generative AI features, Google states directly that its AI surfaces are rooted in its core search ranking and quality systems, and that, in its own words, optimising for generative AI search is optimising for the search experience, and thus still SEO. It treats the AEO and GEO labels as describing the goal of appearing in AI answers rather than a separate method you need to buy.
The same guidance retires several of the tactics the gimmick layer has been selling: it says you do not need a special machine-readable file such as llms.txt, you do not need to chop your content into tiny chunks, and you do not need AI-specific schema or AI-specific rewriting to appear in its AI answers. What it says does matter is the thing SEO always rewarded: genuinely useful, non-commodity content with a real point of view, on a site that is clean, crawlable and indexable. Google went further in mid-2026, publishing separate guidance on evaluating third-party SEO and AEO services and reminding site owners that no external tool can see its ranking signals. That is not a new discipline being born. It is the old one, confirmed by the company whose surface everyone is trying to appear on.
One honest caveat, because the balanced version of this argument has to carry it. Google’s guidance covers Google’s own surfaces. ChatGPT, Perplexity and Claude run their own retrieval and have their own source preferences, which is exactly why their citations diverge so sharply, so there is a genuine off-site, brand-perception layer that classic on-page SEO does not fully address. That is the part of the AEO story with real substance. But notice what that layer actually is: earning credible, consistent third-party coverage so the engines have good sources to draw on. That is PR and digital authority. It is still not a separate, cleanly measurable channel with its own optimisation algorithm. It is the off-site half of good marketing.
SEO and AEO measurement, side by side
The clearest way to see why one is measurable and the other is not is to put the properties that actually matter next to each other. This is not about which is more important. It is about whether each gives you something you can report as a clean, optimisable metric.
| Measurement property | Classic SEO | AEO / AI citation |
|---|---|---|
| The unit being counted | The query: a discrete, repeatable, loggable search term. | The conversation: long, unique, multi-turn, rarely repeated the same way twice. |
| Denominator | Impressions, a real observed population of how often you could appear. | None. The set of possible conversations is effectively infinite and unobservable. |
| Reproducibility | High. The same query returns a stable, comparable result set. | Low. Engines are non-deterministic; the same prompt can return different answers and sources. |
| Where the sample comes from | Real searches, captured in Search Console and analytics. | A prompt set you or the vendor designed, which encodes its author’s assumptions. |
| Cross-engine consistency | One dominant engine to optimise for, with stable signals. | Engines barely agree; around 11 percent domain overlap between major platforms. |
| Attribution to outcome | Traceable, click to session to conversion, with referrer data. | Largely broken; many citations pass no referrer, and cited publishers see under 1 percent referral traffic. |
| The main lever | Useful content plus technical foundations plus authority. | The same, plus third-party mentions, which correlate with citation about 3x more than backlinks. |
| What you can honestly call it | A measurable channel with a reportable return. | A diagnostic signal. Useful to read, not sound to optimise or report as ROI. |
Read down the AEO column and the conclusion writes itself. Every property that makes classic search measurable, a stable unit, a real denominator, reproducibility, a clean sample, attribution, is either missing or broken. What remains is genuinely useful as a diagnostic, and the main lever it points to is the one SEO and PR already own.
So what should you actually do
Care about AI visibility. It is a real and growing surface, and being absent from the answer when a buyer asks about your category is a real cost. Nothing here says ignore it. What it says is: do not buy it as a separate, measurable channel, because it cannot be measured that way without importing the bias of whoever designed the prompt set.
In practice that means a few things. Treat citation-tracking tools as a diagnostic, not a KPI. They are genuinely useful for spotting that you are absent from a conversation you should own, or that a competitor dominates a cluster, or that an engine is describing you inaccurately. They are not a performance dashboard you optimise against or report as return on investment, because the underlying number is a biased sample of a moving target. Track more than one engine if you track at all, because the one you ignore is where your biggest gap probably lives. Put the budget where it actually works: the SEO foundations, the genuinely useful and specific content, and the third-party coverage that earns credible mentions. And be sceptical of any agency selling AEO as a distinct service with its own retainer when the deliverables are an audit, a brief and a roadmap that would be at home in any SEO engagement. Ask them the only question that matters: once you have the citation data, what will you actually do differently? If the honest answer is “improve the content and earn more coverage”, you are buying SEO and PR with a markup for the dashboard.
The discipline is real. The surface is real. The measurement, sold as a clean number you can optimise, is the gimmick. Keep the first two, and do not pay extra for the third.
Key takeaways
- AEO is mostly SEO, and Google says so. Its own 2026 guidance states that optimising for generative AI search is still SEO, on the same ranking and quality systems, with no special files, chunking or schema required.
- It is unmeasurable as a channel, not unmeasurable full stop. Classic search has a stable unit and a real denominator; conversational AI search has neither, so the tools give you a sample, not a measurement.
- The prompt set is the bias. Citation data reflects the assumptions of whoever wrote the prompts, because building them from “what we think people search” imports a keyword mindset into a medium that is not keywords.
- There is no single number. The major engines share only around 11 percent of cited domains, so a blended share-of-voice score averages across systems that fundamentally disagree.
- You usually cannot close the loop. Many AI citations pass no referrer, and heavily-cited publishers see under 1 percent of their traffic from AI platforms, so a citation rarely traces to an outcome.
- The lever is PR. Brand web mentions correlate with AI citations about three times more strongly than backlinks, which means the real driver is credible third-party coverage, a discipline that already exists.
- Use the tools as a diagnostic, not a KPI. They are good for spotting absence, competitor dominance or misrepresentation; they are not a performance dashboard to optimise against or report as ROI.
- The test for any AEO pitch. Ask what you would do differently once you have the citation data. If the answer is “improve content and earn coverage”, you are buying SEO and PR with a markup.
FAQs
Is AEO just SEO with a new name?
Largely, yes, for the on-site work, and Google now says so in its own guidance: optimising for its generative AI features is still SEO, because those features run on the same index and ranking systems as classic search. The same guidance says you can ignore llms.txt files, content chunking and AI-specific schema. The genuine addition is an off-site, brand-perception layer, how credibly and consistently third parties describe you, which feeds the sources engines draw on. But that layer is PR and digital authority, not a new optimisation science. So treat AEO as the SEO and PR you should be doing, pointed at AI surfaces, rather than a separate discipline to buy.
Is AEO worth doing in 2026?
The underlying work is worth doing, because being visible in AI answers is a real and growing surface. But most of what makes you visible in AI answers is good SEO and good PR: useful, specific content on a crawlable site, plus credible third-party coverage. So yes, care about it, but treat it as part of SEO and digital authority rather than a separate channel with its own budget. The thing not worth doing is paying a premium for AEO as a distinct service when the deliverables are the SEO and content work you should already be doing.
Can you actually measure AI search visibility?
Not in the way you can measure classic search, and that is the central problem. Classic search has a stable, countable unit, the query, and a real denominator in impressions, so click-through rates mean something. Conversational AI queries do not repeat, do not match in tidy phrases, and have no observable denominator, and the engines are non-deterministic, so the same question can return different answers and sources. Tracking tools give you a sample of citations over a prompt set someone chose, which is a useful diagnostic but not a clean measurement of a channel.
What do tools like Profound and Scrunch actually do?
They monitor multiple answer engines by running prompts and capturing the responses, then report how often your brand is mentioned or cited, how that compares to competitors as a share of voice, the sentiment of how you are described, and which sources the engines rely on. Some add auditing and content suggestions. That is genuinely useful as a diagnostic. The limits are that the data is a sample over a chosen prompt set rather than a population, the engines barely agree on who to cite, and you usually cannot trace a citation through to a business outcome. Google has also noted that no third-party tool can see its internal ranking data, so any visibility score is a model, not a measurement.
Why is it hard to measure AI citations without bias?
Because the measurement starts with a prompt set that someone designs, and that design encodes assumptions. If you build the prompts from what you think people would search for, you import a keyword mindset into a medium that does not behave like keywords, so the data reflects your assumptions about how people ask rather than how they actually use LLMs conversationally. Change the prompt set and the numbers change. When the result depends that much on the design of the test, you have a reflection of your own assumptions more than an objective metric.
Do different AI engines cite the same sources?
Mostly not, and this is one of the strongest reasons a single AI visibility score misleads. Independent analyses of hundreds of millions of citations have found only around 11 percent of domains cited by both ChatGPT and Perplexity, and even Google’s own AI Overviews and AI Mode share only about 14 percent of cited URLs despite reaching similar answers. The overlap between AI-cited sources and Google’s top organic results has also fallen sharply over two years. The practical implication is that you can be highly visible in one engine and nearly absent in another, so any tool that reads one engine, or blends them into one number, will hide that gap.
Should I stop using AI visibility tools?
No, but use them for what they are good at. As a diagnostic they can show you that you are missing from a conversation you should own, that a competitor dominates a topic cluster, or that an engine is describing your brand inaccurately, all of which are genuinely useful to know. The mistake is treating the output as a clean performance metric you optimise against or report as return on investment, because it is a biased sample of a non-deterministic, unattributable system. Read the signal, act on the SEO and PR it points to, and do not mistake the dashboard for a measurement of a channel.
Last reviewed: June 2026
This article reflects our considered opinion, informed by hands-on work in answer engine optimisation, and is general information rather than a guarantee of any particular outcome in AI search. AI engines, tools and platform guidance change quickly; check the current position before making decisions based on this.
Want to be the answer, not the tenth blue link?
We build SEO on pages you own, show you what ranked and why, then hand it back so it keeps earning without us.
Book an SEO audit →We use cookies to run this site, understand how it is used and improve it. Essential cookies are always on. You can accept everything, reject the rest or choose what runs. Read our cookie policy.
Essential
Needed for the site to work, including security and your cookie choice. These cannot be switched off.
Analytics
Helps us understand which pages people read so we can write better ones. Anonymous and aggregated.
Marketing
Used to measure campaigns and show relevant content across other platforms.