If a business appears regularly in AI recommendations, does that mean it is being positioned equally strongly across different AI platforms?
I tested this by asking Google AI Mode and ChatGPT exactly the same buying question five times each.
The clearest result was this:
| Firm | Google recommended | Google ranked first | ChatGPT recommended | ChatGPT ranked first |
|---|---|---|---|---|
| Linford Grey | 5/5 | 5/5 | 2/5 | 0/5 |
| Agile Accountants | 4/5 | 0/5 | 4/5 | 4/5 |
Google AI Mode placed Linford Grey first in every run.
ChatGPT placed Agile Accountants first in four of its five runs.
But Agile appeared in exactly the same number of responses on both platforms: 4 out of 5.
A measurement based only on whether Agile was recommended would therefore make its performance on Google AI Mode and ChatGPT look almost identical.
It wasn’t.
A brand can have similar recommendation presence on two AI platforms while having very different recommendation prominence.
That is the main finding from this experiment.
Why I ran this test
Previous experiments on this site have shown that AI recommendations can change when:
- an identical prompt is repeated;
- the same customer situation is expressed in different ways;
- the test is conducted at a different time.
That led naturally to another question.
What happens if the prompt stays exactly the same, but the AI platform changes?
If a business monitors its visibility in Google AI Mode, can it assume that its position in ChatGPT will be broadly similar?
Or does each platform need to be measured separately?
The query
I used the same commercially specific accountancy query from earlier experiments:
“I run a small manufacturing company in Birmingham with 25 employees. I need an accountancy firm to handle year-end accounts, corporation tax, payroll and monthly management accounts. I use Xero and would prefer a firm offering fixed monthly fees. Which three Birmingham accountancy firms should I consider, and why?”
The query includes several decision-relevant requirements:
- Birmingham location;
- manufacturing sector;
- 25 employees;
- Xero;
- year-end accounts;
- corporation tax;
- payroll;
- monthly management accounts;
- fixed monthly fees.
Both platforms therefore received exactly the same customer situation.
How I ran the experiment
I ran the query five times in Google AI Mode and five times in ChatGPT.
To keep the tests close together in time, I alternated between the two:
Google 1 → ChatGPT 1
Google 2 → ChatGPT 2
Google 3 → ChatGPT 3
Google 4 → ChatGPT 4
Google 5 → ChatGPT 5
Each run was carried out as a fresh search or conversation.
For this experiment I recorded only:
- which three firms were recommended;
- the order in which they were presented.
I did not analyse citations.
That gave me 10 responses and 30 recommendation positions.
The complete results
| Run | Google AI Mode | ChatGPT |
|---|---|---|
| 1 | Linford Grey / Ballards LLP / Agile Accountants | Agile Accountants / Linford Grey / Tax Care |
| 2 | Linford Grey / Agile Accountants / Companies999 | Agile Accountants / Inform Accounting / Gondal Accountancy |
| 3 | Linford Grey / Morgan Reach / Gondal Accountancy | Agile Accountants / Morgan Reach / Gondal Accountancy |
| 4 | Linford Grey / Gondal Accountancy / Agile Accountants | Gondal Accountancy / Dains / Tax Care |
| 5 | Linford Grey / Ballards LLP / Agile Accountants | Agile Accountants / Prime Accountants / Linford Grey |
The bold business is the first recommendation.
Google AI Mode had one extremely stable leader
Google changed its second and third recommendations across the five runs.
But its first recommendation did not change at all.
| Firm | Google appearances |
|---|---|
| Linford Grey | 5/5 |
| Agile Accountants | 4/5 |
| Ballards LLP | 2/5 |
| Gondal Accountancy | 2/5 |
| Companies999 | 1/5 |
| Morgan Reach | 1/5 |
Linford Grey appeared in every Google AI Mode result.
More importantly:
Linford Grey ranked first in 5 out of 5 Google AI Mode runs.
So Google showed an interesting combination of stability and variation.
The wider shortlist changed, but one business remained anchored at the top.
Agile Accountants also had strong Google visibility, appearing in four of the five searches.
However:
Agile ranked first 0 out of 5 times on Google.
That becomes particularly interesting when we compare it with ChatGPT.
ChatGPT preferred a different business
ChatGPT’s recommendation frequencies were:
| Firm | ChatGPT appearances |
|---|---|
| Agile Accountants | 4/5 |
| Gondal Accountancy | 3/5 |
| Linford Grey | 2/5 |
| Tax Care | 2/5 |
| Inform Accounting | 1/5 |
| Morgan Reach | 1/5 |
| Dains | 1/5 |
| Prime Accountants | 1/5 |
Agile appeared in four of the five responses.
And on every occasion it appeared:
ChatGPT ranked Agile first.
So Agile’s ChatGPT result was:
Recommended: 4/5
Ranked first: 4/5
Compare that with Google:
Recommended: 4/5
Ranked first: 0/5
That is the clearest result in the experiment.
Recommendation presence and recommendation prominence are different
Most AI visibility measurement naturally begins with presence.
Did the business appear?
How often?
That is useful information.
But this experiment suggests that it may not be enough.
Recommendation presence
This asks:
Was the business included in the shortlist?
Recommendation prominence
This asks:
How strongly was the business positioned within that shortlist?
For this experiment I have used first position as a simple measure of prominence.
I have not tested whether customers are actually more likely to click, enquire about or choose the first business rather than the second or third.
So first position should not automatically be treated as a measure of commercial value.
It does, however, show how strongly the AI response presented one business relative to the others.
And Agile demonstrates why that distinction matters.
A presence-only report would say:
Google AI Mode: Agile appeared 4/5.
ChatGPT: Agile appeared 4/5.
Those figures look identical.
A prominence measure reveals:
Google AI Mode: Agile first 0/5.
ChatGPT: Agile first 4/5.
The underlying picture is therefore very different.
Linford Grey showed an equally strong platform difference
Linford Grey produced almost the reverse pattern.
| Metric | Google AI Mode | ChatGPT |
|---|---|---|
| Recommended | 5/5 | 2/5 |
| Ranked first | 5/5 | 0/5 |
Within this sample, Linford Grey was exceptionally prominent in Google AI Mode.
It was much less visible in ChatGPT and was never ranked first there.
Again, this does not establish a permanent platform preference.
Five searches are far too few for that.
But it clearly demonstrates why performance on one platform should not simply be assumed to represent performance on another.
The platforms still agreed on part of the competitive set
Despite the prominence differences, Google AI Mode and ChatGPT did not operate with completely separate groups of firms.
The number of shared recommendations in each paired run was:
| Run | Firms appearing on both platforms |
|---|---|
| 1 | 2/3 |
| 2 | 1/3 |
| 3 | 2/3 |
| 4 | 1/3 |
| 5 | 2/3 |
Every paired test shared at least one recommendation.
Three of the five pairs shared two.
Four businesses appeared at least once on both platforms:
- Linford Grey;
- Agile Accountants;
- Gondal Accountancy;
- Morgan Reach.
So the experiment does not support the claim:
Google AI Mode and ChatGPT recommend completely different businesses.
A more accurate interpretation is:
The platforms partly agreed about the competitive set, while differing substantially in how often and how prominently some businesses appeared.
Some firms appeared on only one platform within this sample
Google AI Mode recommended six different firms across its 15 positions.
ChatGPT recommended eight.
Appeared only in the Google AI Mode sample
- Ballards LLP;
- Companies999.
Appeared only in the ChatGPT sample
- Tax Care;
- Inform Accounting;
- Dains;
- Prime Accountants.
The wording within this sample matters.
Another five searches could easily introduce some of these firms on the other platform.
The experiment does not show that they are permanently platform-specific businesses.
It simply records that their observed visibility differed during these tests.
One combined visibility figure could hide useful information
Across all ten searches, Agile Accountants appeared 8 times.
That is a perfectly valid overall measurement.
But it combines two very different platform results:
Google AI Mode: 4/5 recommended, 0/5 first
ChatGPT: 4/5 recommended, 4/5 first
The problem is therefore not aggregation itself.
An overall figure can still be useful.
The danger is treating the aggregate figure as if it tells the complete story.
For businesses using automated AI visibility trackers, the ability to segment results by platform may be particularly important.
A dashboard might show an overall recommendation Share of Voice.
But you may also want to know:
- how that visibility is distributed across platforms;
- whether competitors dominate particular platforms;
- whether your brand merely appears or is regularly placed first.
Platform-level visibility
This gives us another useful way to think about AI visibility.
Platform-level visibility asks:
How does a brand perform for the same customer requirement on a particular AI platform?
Within these five runs:
- Linford Grey showed very strong Google AI Mode visibility and prominence;
- Agile Accountants showed strong presence on both platforms but much greater prominence on ChatGPT;
- Gondal Accountancy appeared more often on ChatGPT than Google.
Those differences would be lost if the ten searches were treated as one undifferentiated pool.
AI visibility is becoming a multi-dimensional measurement
The experiments on this site are increasingly suggesting that “AI visibility” is not one simple percentage.
It can vary by:
- repetition — the same prompt can produce different answers;
- exact prompt — visibility belongs to the specific test being run;
- customer intent — natural formulations of the same need can produce different patterns;
- platform — Google AI Mode and ChatGPT may treat brands differently;
- prominence — appearing is not necessarily the same as being strongly recommended;
- time — results may change when the experiment is repeated later.
That does not mean AI visibility cannot be measured.
It means the measurement needs context.
What this means for automated AI visibility tracking
The complexity also explains why automated monitoring becomes useful quite quickly.
Imagine tracking:
- multiple customer situations;
- several natural formulations for each;
- Google AI Mode;
- ChatGPT;
- other AI platforms;
- repeated runs;
- changes over weeks and months.
Manual testing becomes difficult to manage at scale.
Automation can handle that workload.
But the measurements still need to be interpreted properly.
A visibility percentage should lead to questions such as:
Which platforms contributed to it?
Which customer situations?
How often was the brand merely present?
How often was it particularly prominent?
The software can collect the observations.
The measurement framework determines what those observations mean.
Limitations
This was a small exploratory experiment.
It involved:
- one customer situation;
- one industry;
- one location;
- one exact prompt;
- two AI platforms;
- five runs per platform;
- ten responses in total.
Five runs are not enough to establish permanent recommendation probabilities.
The test also covered only one short testing period.
Repeating it later could produce very different results.
The experiment therefore does not show that Google AI Mode will always rank Linford Grey first for this query or that ChatGPT will consistently prefer Agile Accountants.
It shows what happened during these ten responses.
The consistency of the first-place recommendations is interesting precisely because it occurred repeatedly within this sample.
What the experiment showed
I began with a simple question:
If I ask Google AI Mode and ChatGPT exactly the same buying question, will they recommend the same businesses?
The answer was:
Partly.
There was meaningful overlap in the businesses considered.
But recommendation frequency and recommendation prominence differed substantially.
The clearest results were:
Google AI Mode ranked Linford Grey first in 5/5 searches.
and:
ChatGPT ranked Agile Accountants first in 4/5 searches.
Perhaps most importantly, Agile appeared in 4/5 searches on both platforms, yet its first-place result was 0/5 on Google and 4/5 on ChatGPT.
That leads to the main lesson from the experiment:
A brand can have similar recommendation presence on two AI platforms while having very different recommendation prominence.
So when measuring AI recommendation visibility, asking simply:
“How often do we appear?”
may not tell the whole story.
It may also be worth asking:
On which platform?
For which customer situation?
How consistently?
And when we appear, how strongly are we being recommended?
Those distinctions can reveal competitive differences that one headline visibility number may hide.
Testing conducted using Google AI Mode and ChatGPT on 10 August 2026. AI-generated recommendations and platform behaviour can change over time.