If a business tracks one carefully chosen AI search prompt, how much does that result tell us about its visibility for the wider customer need?
I tested this by taking one commercially specific accountancy scenario and expressing it five different ways.
The customer did not change.
The location did not change.
The services required did not change.
The Xero requirement did not change.
The preference for fixed monthly fees did not change.
Each prompt described substantially the same buying situation.
I ran each version three times through Google AI Mode, producing 15 responses and 45 recommendation positions.
The clearest result was this:
| Firm | Prompt A | Prompt E |
|---|---|---|
| Linford Grey | 3/3 | 0/3 |
| Gondal Accountancy | 0/3 | 3/3 |
The underlying customer requirement was substantially the same.
Yet one formulation recommended Linford Grey every time and Gondal Accountancy never.
Another formulation produced the opposite result.
At the same time, identical prompts could also generate completely different shortlists from one run to the next.
So this experiment does not show that changing a few words predictably determines which businesses appear.
It shows something more useful:
Visibility for one exact prompt is not necessarily the same as visibility across different natural ways of expressing the same customer need.
Why I ran this experiment
In How to Choose AI Search Prompts That Actually Matter to Your Business, I argued that businesses should build AI visibility testing around genuine customer situations rather than large collections of superficial keyword variations.
I still think that is the right principle.
But it creates another question.
Suppose we identify an important customer situation and create one representative prompt.
Can the resulting visibility measurement then be treated as visibility for the wider customer intent?
Real customers are unlikely to describe identical situations using identical words.
One person might ask:
Which accountants should I consider?
Another might say:
Can you recommend three accountants?
Someone else might ask:
What are three good accountancy firms for a company like mine?
I wanted to see whether Google AI Mode produced similar recommendation patterns when the underlying buying situation remained substantially unchanged but the wording varied naturally.
The customer situation
Every prompt described a prospective customer with the same main characteristics:
- a manufacturing business;
- based in Birmingham;
- employing 25 people;
- using Xero;
- requiring year-end accounts;
- requiring corporation tax support;
- requiring payroll;
- requiring monthly management accounts;
- preferring fixed monthly pricing;
- asking for three accountancy firms to consider.
The purpose was not to test the effect of adding or removing an important buying requirement.
I had already explored that in earlier experiments.
This time I wanted to keep the commercial meaning as consistent as reasonably possible while changing the way it was expressed.
The five prompt formulations
Prompt A
“I run a small manufacturing company in Birmingham with 25 employees. I need an accountancy firm to handle year-end accounts, corporation tax, payroll and monthly management accounts. I use Xero and would prefer a firm offering fixed monthly fees. Which three Birmingham accountancy firms should I consider, and why?”
Prompt B
“Can you recommend three Birmingham accountants for a 25-employee manufacturing company using Xero? We need year-end accounts, corporation tax, payroll and monthly management accounts, preferably for a fixed monthly fee.”
Prompt C
“Which three accountancy firms in Birmingham would be suitable for a manufacturing business with 25 employees that uses Xero and needs payroll, monthly management accounts, corporation tax and year-end accounts on a fixed monthly fee?”
Prompt D
“What are three good Birmingham accountancy firms for a small manufacturer with 25 staff? The company uses Xero and wants payroll, management accounts, year-end accounts and corporation tax included within a fixed monthly package.”
Prompt E
“I am looking for an accountant in Birmingham for a manufacturing company employing 25 people. They must support Xero and provide payroll, monthly management accounts, corporation tax and annual accounts. I would prefer predictable fixed monthly pricing. Which three firms should I look at?”
These formulations were created for the experiment.
They were not collected from real customer search logs or actual sales conversations.
They therefore represent a small set of plausible paraphrases rather than every possible way a customer could express the need.
How I conducted the test
I ran every prompt three times through Google AI Mode.
Rather than completing all three runs of Prompt A before moving to Prompt B, I rotated through them:
Cycle 1: A1 → B1 → C1 → D1 → E1
Cycle 2: A2 → B2 → C2 → D2 → E2
Cycle 3: A3 → B3 → C3 → D3 → E3
Each run began as a fresh search.
I recorded the three recommended firms and their positions.
I did not analyse citations in this experiment. The objective was specifically to compare recommendation visibility.
That produced:
5 formulations × 3 runs × 3 recommendations = 45 recommendation positions.
The complete results
Prompt A
| Run | First | Second | Third |
|---|---|---|---|
| A1 | Linford Grey | Ballards LLP | Companies999 |
| A2 | Companies999 | Linford Grey | Tax Care |
| A3 | Ballards LLP | CASS | Linford Grey |
Linford Grey appeared in 3/3 responses.
Ballards LLP and Companies999 each appeared twice.
Five different firms occupied the nine available positions.
Prompt B
| Run | First | Second | Third |
|---|---|---|---|
| B1 | CASS | Linford Grey | Onyx Accountants |
| B2 | Prime Accountants | CASS | Gondal Accountancy |
| B3 | Linford Grey | Inform Accounting | Morgan Reach |
Seven different firms appeared.
No business appeared in all three runs.
Linford Grey and CASS each appeared twice.
Prompt C
| Run | First | Second | Third |
|---|---|---|---|
| C1 | Dains Accountants | Avonmead Accountants | Linford Grey |
| C2 | Gondal Accountancy | Ken Bell Accounting | BookCheck |
| C3 | Avonmead Accountants | Linford Grey | eCloud Experts |
Again, seven different firms appeared.
C1 and C2 had no recommendations in common at all, despite using identical wording.
C3 then brought both Avonmead and Linford Grey back into the shortlist.
Prompt D
| Run | First | Second | Third |
|---|---|---|---|
| D1 | Inform Accounting | Gondal Accountancy | Linford Grey |
| D2 | Linford Grey | Inform Accounting | Gondal Accountancy |
| D3 | Linford Grey | Tax Care | Inform Accounting |
Prompt D produced the tightest competitive group.
Only four different firms appeared across its nine positions.
Two businesses appeared every time:
Linford Grey — 3/3
Inform Accounting — 3/3
Gondal Accountancy appeared twice.
Prompt E
| Run | First | Second | Third |
|---|---|---|---|
| E1 | Gondal Accountancy | eCloud Experts | Apex Accountants |
| E2 | Audit Consulting Group | Apex Accountants | Gondal Accountancy |
| E3 | Audit Consulting Group | Agile Accountants | Gondal Accountancy |
Gondal Accountancy appeared in 3/3 responses.
Audit Consulting Group and Apex Accountants each appeared twice.
Linford Grey appeared 0/3.
Some formulations produced much narrower recommendation pools
The five formulations did not show the same level of repeatability.
| Prompt | Different firms across 9 recommendation positions |
|---|---|
| A | 5 |
| B | 7 |
| C | 7 |
| D | 4 |
| E | 5 |
Prompt D repeatedly drew from a relatively concentrated group.
Prompts B and C produced a wider range of businesses.
So variation across the formulations was not limited to individual company names.
The apparent concentration of the competitive pool also differed.
Linford Grey shows why one exact prompt can be misleading
Linford Grey was the most frequently recommended firm in the experiment.
But its appearances were distributed very unevenly.
| Prompt | Linford Grey |
|---|---|
| A | 3/3 |
| B | 2/3 |
| C | 2/3 |
| D | 3/3 |
| E | 0/3 |
| All searches | 10/15 |
If I had tracked only Prompt A, the result would have been:
Linford Grey appeared in every observed run.
If I had tracked only Prompt E:
Linford Grey did not appear at all.
Both observations would have been correct.
But neither would have described its visibility across the wider set of plausible formulations used in the experiment.
Across all 15 searches, Linford appeared 10 times.
Gondal Accountancy produced almost the reverse pattern
Gondal Accountancy appeared seven times overall.
Its distribution was:
| Prompt | Gondal Accountancy |
|---|---|
| A | 0/3 |
| B | 1/3 |
| C | 1/3 |
| D | 2/3 |
| E | 3/3 |
| All searches | 7/15 |
The contrast between A and E is particularly striking.
Prompt A gave:
Linford Grey — 3/3
Gondal Accountancy — 0/3
Prompt E gave:
Linford Grey — 0/3
Gondal Accountancy — 3/3
The prompts did not describe fundamentally different customers.
Yet the observed recommendation patterns were completely reversed.
Three runs per formulation are far too few to claim that particular wording caused those outcomes.
But they are enough to demonstrate that the observed visibility of the two firms was not evenly distributed across the five formulations.
Inform Accounting showed another concentrated pattern
Inform Accounting appeared four times in total.
Three of those appearances came from Prompt D.
| Prompt | Inform Accounting |
|---|---|
| A | 0/3 |
| B | 1/3 |
| C | 0/3 |
| D | 3/3 |
| E | 0/3 |
So the firm appeared in every observed run of Prompt D but only once across the other 12 searches.
Again, this does not establish a causal relationship between wording and recommendation.
It does show why a result from one monitored prompt should not automatically be treated as representative of every natural formulation of the same customer need.
Identical prompts could also produce very different results
This is an important limitation on how the experiment should be interpreted.
Changing the wording was not necessary for the recommendations to change.
Prompt C provides the clearest example.
C1 returned:
- Dains Accountants;
- Avonmead Accountants;
- Linford Grey.
C2 returned:
- Gondal Accountancy;
- Ken Bell Accounting;
- BookCheck.
There was no overlap at all.
Nothing in the prompt had changed.
This is consistent with my earlier experiment, I Ran the Same Google AI Mode Product Search 10 Times—Here’s How the Recommendations Changed, where identical searches also produced changing shortlists.
So the experiment does not demonstrate:
Change the wording and Google will predictably change the recommendation.
Instead, I observed two overlapping forms of variation.
Repeated-response variation
An identical prompt can produce different businesses on different runs.
Variation across formulations
The observed recommendation patterns can also differ across several natural ways of expressing substantially the same customer requirement.
Separating those two effects is important.
18 different firms appeared
Across the 45 recommendation positions, 18 different accountancy firms appeared.
The most frequent were:
| Firm | Appearances |
|---|---|
| Linford Grey | 10 |
| Gondal Accountancy | 7 |
| Inform Accounting | 4 |
| CASS | 3 |
| Ballards LLP | 2 |
| Companies999 | 2 |
| Tax Care | 2 |
| Avonmead Accountants | 2 |
| eCloud Experts | 2 |
| Apex Accountants | 2 |
| Audit Consulting Group | 2 |
Another seven firms appeared once.
So the recommendations were highly variable.
But they were not simply a random collection of one-off names.
Linford Grey and Gondal Accountancy appeared much more frequently than most competitors.
The important point is that even their stronger visibility was distributed unevenly across the five formulations.
Exact-prompt visibility and intent-level visibility
I think this is the most useful distinction to emerge from the experiment.
Exact-prompt visibility
This asks:
How frequently does the business appear when one precise saved prompt is tested?
That is a legitimate measurement.
A visibility tracker can repeat that prompt over time and record whether a brand appears.
But the result belongs to that prompt.
It should not automatically be interpreted as visibility for every customer expressing the same underlying need.
Intent-level visibility
For the purposes of this experiment, I use intent-level visibility to mean:
Visibility across a small set of plausible formulations representing substantially the same customer situation.
It does not mean that five prompts capture every possible way a real customer could express the intent.
But it provides a broader view than one exact formulation alone.
Linford Grey demonstrates the difference.
Its exact-prompt results ranged from:
3/3 appearances
to:
0/3 appearances
depending on which formulation was monitored.
Across the complete 15-search set, it appeared 10 times.
Those measurements answer different questions.
This does not mean tracking dozens of trivial rewrites
The conclusion is not that businesses should monitor 50 tiny wording changes such as:
- best accountant;
- good accountant;
- recommended accountant;
- top accountant;
- accountant to consider.
That risks recreating traditional keyword tracking with AI prompts.
The customer situation should remain the starting point.
But this experiment suggests that one exact prompt may also be too narrow to represent an entire important customer intent.
A more practical approach could be:
- Identify a commercially meaningful customer situation.
- Create a small number of genuinely natural formulations.
- Keep those formulations unchanged when measuring trends.
- Record exact-prompt performance.
- Also review visibility across the group as a broader indication of intent-level visibility.
The right number of formulations is not established by this experiment.
But it does suggest that, for an important customer situation, one may not always be enough.
Why automated tracking still matters
This experiment required only 15 searches.
That was manageable manually.
Now imagine monitoring:
- 20 customer situations;
- three or five formulations for each;
- several AI platforms;
- multiple repetitions;
- changes every week or month.
The volume becomes substantial very quickly.
That is where automated AI visibility trackers become useful.
But automation solves the measurement workload.
It does not remove the need to decide what the measurements represent.
A tracker might accurately report:
Brand X appeared in 3/3 runs of Prompt A.
That result can be perfectly correct.
The mistake would be silently turning it into:
Brand X has complete visibility for customers with this general need.
Our experiment shows why those statements are not equivalent.
The same exact prompt also changed across testing periods
There was one final result worth noting.
Prompt A was the same wording I had used during an earlier ten-run accountancy experiment.
In that earlier set:
Linford Grey appeared 0 times out of 10.
During this experiment:
Linford Grey appeared in all 3 Prompt A runs.
So Prompt A itself did not produce a stable long-term result.
That means the patterns observed here cannot be explained simply by different wording.
The testing period also appears to matter.
The broader lesson is:
An AI visibility result is an observation made under particular testing conditions. It should not be mistaken for a permanent ranking attached to a prompt.
Limitations
This was a small exploratory experiment.
It used:
- one customer situation;
- one industry;
- one location;
- one AI platform;
- five researcher-created formulations;
- three repetitions of each;
- 15 searches in total.
Three repetitions are not enough to estimate the true underlying probability that a business will appear for a particular formulation.
The five prompts also preserved the same overall commercial situation rather than being linguistically identical.
Small changes in wording, emphasis and sentence structure may influence how the request is interpreted.
That possibility is part of what the experiment was intended to explore, but the results do not establish causation.
The experiment therefore demonstrates observed recommendation patterns, not permanent or statistically proven visibility rates.
What the experiment showed
I began with the question:
If two customers describe substantially the same buying situation differently, will businesses have similar AI visibility?
The answer from these 15 searches was:
Not necessarily.
Across five plausible formulations:
- 18 different firms appeared;
- identical prompts sometimes produced completely different shortlists;
- some formulations produced much narrower competitive pools than others;
- Linford Grey appeared 3/3 under two formulations but 0/3 under another;
- Gondal Accountancy appeared 0/3 under Prompt A but 3/3 under Prompt E;
- Inform Accounting appeared 3/3 under Prompt D but only once across the other 12 searches.
The experiment therefore suggests an important distinction:
Exact-prompt visibility tells you what happened for one fixed test. Intent-level visibility asks whether that visibility survives several plausible ways of expressing substantially the same customer need.
Both can be useful.
They should not be confused.
So when constructing an AI visibility monitoring programme, the question may not only be:
Which customer situations should we track?
It may also be:
Is one prompt enough to represent each important situation?
This experiment does not provide a definitive answer.
It does provide a good reason to test the assumption.
Testing conducted using Google AI Mode on 10 August 2026. AI-generated recommendations and Google AI Mode behaviour can change over time.