I asked Google AI Mode a very simple question:
“measure ai brand visibility”
Then I asked it again.
And again.
Five times in total, using exactly the same wording.
The answers were remarkably consistent. Google AI Mode repeatedly referred to the same new industry framework, recommended many of the same metrics, and suggested broadly similar ways for businesses to test their visibility across AI platforms.
But all five answers also omitted the same important point.
Not one told me to repeat an individual prompt several times to measure how much the AI response naturally varies.
That matters because repeated responses are an important part of the measurement framework AI Mode itself was relying on.
And there is a certain irony here.
I only discovered the weakness in AI Mode’s advice because I did exactly what its answers failed to recommend:
I repeated the same prompt.
A Very Simple Five-Run Test
The test was deliberately uncomplicated.
I entered exactly the same prompt into Google AI Mode five times:
measure ai brand visibility
I did not change the wording between runs.
I was originally interested in seeing what advice Google AI Mode would give someone looking for a broad introduction to measuring AI visibility.
The repeated searches soon became an experiment of their own.
What would remain consistent?
What would change?
And, importantly, would AI Mode recognise that AI-generated answers themselves can vary from one identical query to the next?
AI Mode Quickly Settled on the IAB’s 4 Ps
One of the most interesting aspects of the answers was their reliance on a very recent framework.
On 3 August 2026, the Interactive Advertising Bureau published Measuring Visibility in the AI Era, introducing a standardised measurement framework built around what it calls the 4 Ps of AI Visibility.
Only a few days later, all five of my Google AI Mode responses used that framework.
The four categories are:
- Presence — whether the brand appears at all;
- Prominence — how prominently it appears;
- Portrayal — how accurately and favourably it is described;
- Persuasion — whether the response encourages further consideration or action.
AI Mode also repeatedly recommended tracking things such as:
- brand mentions;
- citations;
- Share of Voice;
- recommendation position;
- sentiment;
- AI referral traffic.
So at the headline level, its answer was surprisingly stable.
What Stayed Consistent Across Five Runs?
Here is a simplified comparison of the five responses.
| Recommendation | Run 1 | Run 2 | Run 3 | Run 4 | Run 5 |
|---|---|---|---|---|---|
| IAB 4 Ps referenced | Yes | Yes | Yes | Yes | Yes |
| Manual testing recommended | Yes | Yes | Yes | Yes | Yes |
| Test across multiple AI platforms | Yes | Yes | Yes | Yes | Yes |
| Track mentions/citations | Yes | Yes | Yes | Yes | Yes |
| Share of Voice discussed | Yes | Yes | Yes | Yes | Yes |
| HubSpot mentioned | Yes | Yes | Yes | Yes | Yes |
| Adobe mentioned | Yes | Yes | Yes | Yes | Yes |
| Similarweb mentioned | Yes | Yes | Yes | Yes | Yes |
| Suggested prompt quantity | 20–40 | 20–30 | 20–30 | 20–30 | 20–40 |
| Repeat individual prompts several times | No | No | No | No | No |
That final row became the most interesting result of the test.
What Every Answer Missed
One of the fundamental difficulties with measuring AI visibility is that AI-generated answers are not fixed.
Run exactly the same prompt twice and the model may return:
- different brands;
- different recommendation orders;
- different citations;
- different wording;
- different emphasis.
The IAB framework explicitly recognises this.
Its methodology treats AI visibility as something that needs to be sampled repeatedly rather than judged from a single response.
That distinction is crucial.
Imagine a business creates 20 prompts and tests each one once.
Its brand appears in eight responses.
It records:
40% visibility.
A month later, it repeats the exercise.
This time it appears in 11:
55% visibility.
At first glance, that looks like a 15-percentage-point improvement.
But is it?
Possibly.
Or perhaps the second batch simply produced more favourable AI responses.
Without knowing how much those same prompts vary when repeatedly run under similar conditions, it is difficult to distinguish a real change in visibility from normal response variation.
A Fixed Prompt Set Solves Only Half the Problem
AI Mode repeatedly recommended creating a fixed or “frozen” prompt set.
That is sensible.
If you want to compare visibility over time, constantly changing the questions would make those comparisons much less meaningful.
But there are really two different consistency problems.
Prompt consistency
Use the same defined prompts so that results from different testing periods can reasonably be compared.
Response sampling
Run the same individual prompts enough times to understand how much the AI response itself naturally varies.
Across five answers, AI Mode repeatedly addressed the first issue.
It never properly addressed the second.
Where AI Mode’s Manual Audit Falls Short
The gap becomes clearer when we compare the advice directly with the IAB framework.
| AI Mode repeatedly suggested | IAB framework |
|---|---|
| Test around 20–30 or 20–40 prompts | Fewer than 50 queries should be treated as exploratory |
| Run prompts across several AI platforms | Multiple responses per query are needed to capture normal variation |
| Keep prompts fixed over time | Also establish normal run-to-run variability |
| Calculate mention or Share of Voice percentages | Interpret changes cautiously and account for measurement variability |
This does not mean AI Mode’s suggested manual audit is useless.
Far from it.
A business manually testing 20 carefully chosen prompts could learn a great deal.
It might discover:
- which types of questions trigger its brand;
- which competitors appear instead;
- which websites are cited;
- whether its own site appears as a source;
- how prominently the brand is positioned;
- how it is described.
That can be extremely useful exploratory work.
The problem comes when an exploratory test is presented as though it were a precise measure of overall AI visibility.
There is an important difference between saying:
“We tested 25 commercially important prompts and appeared in seven.”
and saying:
“Our AI visibility is 28%.”
The first describes exactly what happened.
The second risks making a limited test sound far more precise than it really is.
The Answers Were Stable, but the Supporting Details Were Not
The five runs also revealed another useful pattern.
The central synthesis remained very stable while some of the supporting details changed.
Three tools appeared in all five answers:
- HubSpot AI Search Grader;
- Adobe Brand Visibility;
- Similarweb AI Brand Visibility.
Akamai appeared in some runs but not others.
The suggested number of prompts moved between 20–30 and 20–40.
One response specifically suggested monthly testing.
The scoring suggestions were not identical.
And many of the supporting citation sources changed between runs.
Yet the overall answer remained recognisably the same.
This suggests an important distinction when testing AI visibility:
answer stability is not the same thing as citation stability.
An AI system may keep giving essentially the same recommendation while changing some of the sources used to support it.
That has obvious implications for businesses checking whether their site appears as an AI citation.
Appearing once does not necessarily mean you will appear next time.
Disappearing once does not necessarily mean you have somehow “lost” your visibility either.
Measurement Is Not the Same as Optimisation
Another interesting feature of the responses was the tendency to blur AI visibility measurement with terms such as AEO and GEO.
One response even said businesses needed to move from traditional SEO to Answer Engine Optimisation.
But that was not what I asked.
The prompt was simply:
measure ai brand visibility
Measurement and optimisation are related, but they are not the same thing.
Measurement asks:
- Are we appearing?
- For which prompts?
- How often?
- Against which competitors?
- Are we cited?
- Where do we appear in the answer?
- How are we described?
- Does that change over time?
Optimisation asks a different question:
What can we do to change those results?
It makes sense to understand the first before jumping to the second.
AI visibility testing tells you what is happening.
Optimisation tries to influence what happens next.
AI Mode Wasn’t Simply Wrong
The interesting conclusion from this experiment is not that Google AI Mode gave bad answers.
In many respects, the responses were impressive.
All five quickly identified the same recently published industry framework.
They consistently surfaced useful concepts such as Share of Voice, citations and competitor comparison.
They recognised the value of fixed prompt sets and testing across multiple AI platforms.
The more interesting issue was subtler.
AI Mode captured the headline ideas extremely consistently while repeatedly dropping an important methodological qualification.
That is potentially a wider characteristic of AI-generated synthesis.
An answer can preserve the main conclusion while losing some of the nuance or limitations contained in the underlying material.
What a Sensible Manual AI Visibility Test Can Still Do
For a small business or someone beginning to explore AI visibility, manual testing can still be very useful.
A sensible exploratory process might look like this:
- Choose a defined group of commercially meaningful prompts.
- Keep the wording of those prompts fixed.
- Test them across the AI platforms relevant to your audience.
- Run individual prompts more than once rather than treating one response as definitive.
- Record raw observations such as mentions, citations, recommendation position and competitors.
- Treat a small prompt set as exploratory rather than as a precise measure of overall visibility.
- Repeat the process over time using the same methodology.
You do not necessarily need to invent a complicated visibility score.
A spreadsheet showing exactly what happened can often be more useful than compressing limited observations into an impressive-looking percentage.
As the number of prompts, platforms and repeated runs increases, however, manual testing quickly becomes time-consuming.
That is where automated AI visibility monitoring starts to make much more sense.
But automation does not remove the underlying measurement problem.
It simply makes it possible to perform much more sampling.
A Note on This Test
Five runs are not intended to prove how Google AI Mode will answer this query every time.
They simply show what happened across five repeated tests using exactly the same prompt.
That is enough to compare what remained stable, what changed and what was repeatedly omitted.
It is not enough to claim that Google AI Mode always behaves this way.
That distinction is important, particularly in an article about measurement.
The Most Interesting Part of the Experiment
Google AI Mode told me five times how to measure a system whose answers can vary from one run to the next.
Yet none of those five answers told me to repeat an individual prompt enough times to measure that variation.
I only discovered the problem because I did exactly that.
I asked the same question five times.
And that may be the simplest demonstration of why a single AI visibility check should never be mistaken for AI visibility measurement.