When we started testing AI visibility, the question seemed fairly simple:
Does the business appear in the AI-generated answer?
Our experiments increasingly suggest that this is only the first layer.
A useful assessment also needs to consider which query produced the appearance, how consistently it occurs, whether the business is mentioned or actually recommended, how it is portrayed, what evidence supports the answer and whether that visibility has any commercial value.
That has changed how we think about AI visibility testing.
We are no longer looking only for presence.
We are trying to understand the relationship between:
- the customer’s situation;
- the query being tested;
- the competing businesses or information sources;
- the evidence available;
- the way the AI interprets that evidence;
- and the final answer shown to the user.
This page brings together the main lessons emerging from our experiments so far.
These are not claimed Google ranking factors, and we cannot see the internal retrieval or ranking processes used by AI search systems.
They are practical observations and working models based on the tests we have carried out.
If you are primarily looking for instructions on testing your own business, start with our practical AI visibility testing guidance.
This page takes a different route: what are all these experiments beginning to tell us about AI visibility itself?
1. AI Visibility Is More Than Appearing
The most basic AI visibility measurement is presence:
Did the business appear?
That is useful.
But it does not tell us enough.
A business might be:
- mentioned in passing;
- included in a shortlist;
- recommended as the preferred option;
- cited as an information source;
- presented prominently;
- buried near the bottom of an answer;
- described accurately;
- or described in a way that gives the wrong impression.
Those outcomes are not equivalent.
A website could be cited because it contains useful information while completely different businesses are recommended to the customer.
Conversely, a business could be strongly recommended without its own website being cited.
We explored this directly in Does Being Cited by AI Mean Your Business Is Being Recommended?.
This has led us to think about AI visibility as having at least two dimensions.
Visibility quantity
How often are we present?
That can include:
- mentions;
- citations;
- recommendation frequency;
- Share of Voice;
- and position within the answer.
Visibility quality
What happens when we are present?
That introduces different questions:
- Are we being recommended or merely mentioned?
- Are we described accurately?
- Does the AI understand what we actually do?
- Are we being associated with the right customer situations?
- Are important qualifications preserved?
- Does a cited page really support the claim being made?
Two businesses could therefore achieve the same headline visibility percentage while having very different commercial outcomes.
Simply counting appearances is only the beginning.
2. One AI Answer Is an Observation, Not a Measurement
One of our earliest experiments involved running exactly the same Google AI Mode product search ten times.
The recommendations changed considerably.
Six different products appeared across the ten runs. Four different products occupied first place. Only one product appeared every time.
The full experiment is documented in I Ran the Same Google AI Mode Product Search 10 Times—Here’s How the Recommendations Changed.
This exposed a basic problem with one-off AI visibility checks.
If your business appears first once, that tells you what happened in that particular response.
It does not establish that your business consistently occupies first position.
Likewise, failing to appear once does not prove that your business is generally absent.
Repeated testing helps us distinguish between:
consistent visibility, occasional visibility and apparent absence.
But repetition creates another question:
How many times should the same search be repeated?
Our later experiments suggest that there is no universal number.
The correct testing method depends on what you are trying to establish.
A business owner performing a quick manual sense-check is asking a different question from a researcher trying to estimate a true probability with statistical confidence.
For a simple manual repeatability check, we often use five identical runs as a practical starting point.
That does not make five searches statistically significant.
The purpose is much narrower:
Can we see obvious variation when the same query is repeated under broadly similar conditions?
We explain the distinction in How Many Times Should You Repeat an AI Search When Testing Brand Visibility?.
The broader lesson is:
Choose the testing method according to what you are trying to learn.
Do not choose an arbitrary number of searches first and assume that number makes the test meaningful.
3. Test Customer Situations, Not Just Broad Queries
Repeated testing is only useful if you are testing queries that actually matter.
This has become one of the strongest themes running through our experiments.
Imagine an accountancy firm testing:
Recommend an accountant for a small business.
The search is relevant.
But it tells us very little about the customer.
Compare it with someone looking for an accountant:
- in a particular location;
- for a manufacturing company;
- with 25 employees;
- using Xero;
- requiring payroll;
- wanting monthly management accounts;
- and preferring fixed monthly fees.
Those details do more than make the query longer.
They describe the circumstances that determine whether one business is a better fit than another.
In several of our experiments, adding this kind of decision-relevant detail substantially changed the businesses recommended.
That leads to an important distinction.
The question is not simply:
How visible is our business in AI search?
A commercially stronger question is:
How visible are we when potential customers describe situations our business is particularly well suited to solve?
This is why we increasingly favour building AI visibility tests around customer-need themes rather than trying to discover every possible query someone might type.
A customer need might involve:
- a particular problem;
- industry;
- location;
- budget;
- urgency;
- integration;
- previous experience;
- or combination of requirements.
There may be dozens of ways to phrase the same underlying need.
You do not necessarily need to monitor them all.
Our article You Don’t Need to Discover Every AI Search Query: Test Customer-Need Themes Instead explains this approach in more detail.
For businesses creating an actual testing set, How to Build an AI Visibility Query Set Around Real Customer Situations provides the practical method.
There is an important measurement distinction here too.
When exploring, it can be useful to vary the wording and circumstances to understand how recommendations change.
When monitoring, the chosen prompts should usually remain consistent so that changes over time are easier to interpret.
The objective is therefore not to chase every wording.
It is to identify the customer situations worth measuring.
4. Relevance Gets You Into the Competition — It Does Not Guarantee Selection
Our recent experiments have pushed us towards another important distinction.
A webpage can be:
- indexed;
- directly relevant;
- detailed;
- genuinely useful;
- and specifically written to answer the question being tested;
and still not be selected.
We tested this unusually directly.
First, we established what Google AI Mode said about a particular question.
We then created a webpage specifically addressing that issue.
Google Search Console confirmed that the page had been indexed.
We repeated the original broad query another ten times.
The new page was cited:
0 times out of 10.
We then changed the user’s information need progressively.
Eventually, when we asked AI Mode to find a resource explaining the particular distinction our article had been created around, it found and cited the exact page.
The complete experiment is documented in We Built a Page for an AI Search Query — Then Tested What It Took for AI Mode to Cite It.
We should be careful about what this proves.
It does not establish that making a query more specific automatically increases the chance of a page being cited.
Nor does it reveal Google’s internal ranking system.
But it demonstrates something important:
Being relevant and available for selection does not guarantee being selected.
We find it useful to think about this as a competitive environment.
For any information need, there may be many businesses or webpages capable of contributing a plausible answer.
As the user’s requirements change, that competitive environment may change too.
Again, this is a working model for interpreting our experiments, not a description of Google’s internal architecture.
Our current thinking can be summarised as:
Relevant → Competitive → Distinctive → Supported
A page first needs to be relevant.
But if many strong alternatives are equally relevant, something else may determine why one source becomes useful.
That might include:
- closer alignment with the particular requirement;
- specialist information;
- first-party evidence;
- original testing;
- a useful distinction;
- stronger supporting evidence;
- or information that competing sources do not provide.
This led us to another principle:
If your information is interchangeable, your webpage may be interchangeable as a source.
That does not mean websites should deliberately contradict established information.
Distinctiveness is not the same as being contrarian.
Suppose most reliable sources give the same answer because that answer is appropriate for most people.
A useful page does not need to argue that everyone else is wrong.
It might instead explain:
The usual answer is correct in most circumstances, but this particular condition changes the situation.
That is evidence-backed distinctiveness rather than disagreement for its own sake.
We explore that tension further in Consensus vs Distinctiveness: Can a Minority View Compete in AI Search?.
The broader idea is examined in Why Being Relevant Is Not Enough for AI Visibility.
5. AI Search Interprets Information — It Does Not Merely Retrieve It
Another lesson has emerged from looking closely at the answers themselves.
Sometimes an AI response is weak because important information is missing.
But sometimes the necessary facts already exist.
The weakness lies in the conclusion drawn from them.
We call this an Interpretation Gap:
An Interpretation Gap occurs when the relevant facts are available, but the conclusion drawn from them misses an important distinction, condition or exception.
Our repeated-search methodology experiment gave us one example.
The underlying information about sample sizes, variability, repeatability and statistical reliability existed.
But different testing objectives were being compressed into something close to one general recommendation.
When the user’s actual purpose was made explicit, the answer changed.
The topic had barely changed.
The circumstances and objective had.
This is different from what we call a Synthesis Gap.
Synthesis Gap
The necessary information is scattered across several sources and no single resource connects it effectively.
Interpretation Gap
The information is substantially available, but the conclusion drawn from it is too broad, too simple or misses an important qualification.
The distinction matters because the content opportunity is different.
Sometimes a useful page needs to bring scattered facts together.
Sometimes it needs to explain why an existing conclusion requires a distinction.
We explore the concept in The Interpretation Gap: When the Facts Exist but the AI Conclusion Is Too Simple.
There is also a practical testing method in How to Find an Interpretation Gap in an AI Answer.
The basic sequence is:
Identify the conclusion → inspect the evidence → uncover the assumption → identify the circumstance where that assumption stops holding → test it.
Citations Need Interpreting Too
The same issue appears when AI systems cite sources.
Eventually, Google AI Mode cited our page about how many times an AI visibility query should be repeated.
But the resulting answer also introduced a figure of 50+ queries for a more formal statistical exercise.
Our page had not established a universal 50+ rule.
The citation was genuine.
But that did not mean our article was responsible for every part of the surrounding answer.
AI-generated answers can combine information from several sources with additional interpretation and synthesis.
That means:
A citation shows that a page was associated with the answer. It does not prove that every nearby statement came from that page or is supported by it.
When a website is cited, we therefore recommend asking three questions:
Was our page cited?
Then:
What claim appears to rely on it?
Then:
Does our page actually support that claim?
That reveals much more than simply recording:
Citation: Yes
The experiment is examined in detail in Being Cited by AI Does Not Mean It Used Your Page Accurately.
The wider lesson is that AI visibility has to be assessed after the AI has interpreted and synthesised the information, not merely after a URL appears in the citations.
6. Commercial Value Matters More Than the Headline Visibility Score
All of these findings eventually return us to the business question.
Suppose your company appears in 80% of AI responses for a particular query.
That sounds impressive.
But several questions still remain.
What was the customer actually asking?
Was the query commercially relevant?
Was your company being recommended or simply mentioned?
Was it described accurately?
Was the customer shown a useful route towards your business?
A high visibility percentage for an irrelevant query may have little commercial value.
A lower percentage for a highly specific customer situation close to a purchasing decision could be much more important.
This is why we developed the RAMP Framework: Measuring the Commercial Value of AI Visibility.
RAMP considers four stages:
R — Relevance
Does this query represent a customer situation worth being visible for?
A — Appearance
Does the business actually appear, and how consistently?
M — Mention Quality
How is the business presented?
P — Pathway
Can the customer realistically move from the AI answer towards the business?
The order matters.
Commercial relevance comes before the visibility percentage.
The framework therefore brings many of the lessons from our experiments together.
AI visibility is not simply about being present.
It is about being present in the right situation, in the right way, with a meaningful opportunity for the visibility to matter.
What This Means for Businesses
We cannot control how an AI system constructs every response.
We cannot guarantee that a particular webpage will be selected.
We cannot see Google’s internal candidate selection or ranking processes.
And our experiments do not establish that changing one website factor causes an AI recommendation.
But businesses do control something important:
The information they make available about themselves.
A website can clearly explain:
- who the business helps;
- what services it actually provides;
- the situations where it is particularly suitable;
- industries served;
- locations covered;
- prices or budget ranges where appropriate;
- integrations;
- turnaround times;
- important limitations;
- eligibility requirements;
- and the evidence supporting important claims.
Compare a general statement such as:
We work with manufacturers.
with:
We provide monthly management accounts for manufacturing businesses using Xero, including reporting around stock, margins, payroll and cash flow, on a fixed monthly fee.
The second statement gives the customer far more information with which to judge suitability.
It also gives an AI system more concrete evidence about the circumstances in which the business might be relevant.
That does not guarantee visibility.
But it reduces ambiguity about what the business actually offers.
And that is useful regardless of whether the reader is a human or an AI system.
Where Our Thinking Stands Today
Our experiments started with presence:
Did the business appear?
They have progressively pushed us towards:
context, consistency, selection, distinctiveness, evidence, interpretation, portrayal and commercial value.
The question we now find more useful is:
Across the customer situations that matter, how consistently does the business appear, how is it represented, why might it be selected rather than the alternatives, and does that visibility create a meaningful commercial opportunity?
That is considerably harder to measure than a simple mention count.
But it also tells us much more.
Some future experiments may strengthen the working ideas described on this page.
Others may challenge them.
That is exactly what we want.
AI Visibility Testing is not intended to create another collection of universal rules about how AI search supposedly works.
The aim is simpler:
Test what actually happens, separate observation from assumption, follow the evidence and gradually build a better understanding of what AI visibility means for a real business.