I Ran the Same Google AI Mode Product Search 10 Times—Here’s How the Recommendations Changed

If a business wants to know whether its product appears in Google AI Mode, the obvious approach is to conduct a relevant search and look at the answer.

But is one search enough?

I tested this by submitting exactly the same product-selection query to Google AI Mode ten times. Although some products appeared much more consistently than others, both the shortlist and the order of the recommendations changed.

The results illustrate why finding your product once—or failing to find it once—does not provide a reliable measurement of its overall AI visibility.

The query I tested

I used a detailed query based on a plausible software-buying scenario:

“I run a small UK marketing agency with eight employees. We need project management software that includes time tracking, allows clients to view project progress, integrates with Xero, and costs no more than £100 per month. Which three options should we shortlist, and why? Please present the 3 option shortlist as a table just with the recommended names.”

This is considerably more specific than a traditional search such as “best project management software.”

It describes:

  • the type of business;
  • the number of employees;
  • the required features;
  • an existing accounting platform;
  • a fixed monthly budget;
  • the number of recommendations required.

The final instruction asked AI Mode to provide the three names in a simple table. This made the answers easier to capture and compare.

How I conducted the test

I ran all ten searches on 3 August 2026 over approximately one hour.

For every run:

  • I used the same Google account.
  • Search personalisation was turned off.
  • I started a fresh Google search.
  • I pasted the exact same prompt.
  • I selected Google AI Mode.
  • I did not ask any follow-up questions.
  • I captured a screenshot of the three-product shortlist.

Google did not offer a choice between different AI Mode models during these tests.

This was not intended to determine which software was genuinely the best choice. I was testing which products appeared and where they were placed when the same query was repeated.

Run 1

Results of first run repeating same query and logging recommendations

Run 1 returned:

  1. ClickUp
  2. Teamwork
  3. Xero Projects

If I had stopped after this search, I might have concluded that ClickUp had the strongest visibility for the query.

The subsequent runs provided a very different picture.

The complete results

RunFirst recommendationSecond recommendationThird recommendation
1ClickUpTeamworkXero Projects
2Xero ProjectsTeamworkClickUp
3TeamworkProductive.ioPaymo
4TeamworkClickUpProductive.io
5AvazaTeamworkClickUp
6ClickUpTeamworkXero Projects
7TeamworkPaymoClickUp
8Xero ProjectsTeamworkAvaza
9TeamworkProductive.ioPaymo
10TeamworkClickUpXero Projects

The ten responses created 30 recommendation positions, occupied by six different products.

How frequently did each product appear?

ProductAppearancesAppearance rateFirst-place appearances
Teamwork10 out of 10100%5
ClickUp7 out of 1070%2
Xero Projects5 out of 1050%2
Productive.io3 out of 1030%0
Paymo3 out of 1030%0
Avaza2 out of 1020%1

Teamwork had the strongest visibility in this experiment. It appeared in every response, ranked first five times and never appeared below second place.

ClickUp appeared in seven responses, while Xero Projects appeared in five. Productive and Paymo each appeared three times, and Avaza appeared twice.

These patterns could not have been identified from one search.

Avaza shows the danger of relying on one result

Run 5 produced one of the most revealing results.

Results of first run repeating same query and logging recommendations

Run 5 returned:

  1. Avaza
  2. Teamwork
  3. ClickUp

Avaza appeared at the top of this response. However, it was absent from eight of the ten runs.

If Avaza conducted only Run 5, it could conclude that it had excellent visibility and was the leading recommendation.

If it conducted one of the eight searches in which it was absent, it could reach the opposite conclusion.

Neither observation would reveal its wider pattern across the experiment: Avaza appeared in 20% of the responses and ranked first once.

The results varied, but they were not completely random

Two ordered shortlists appeared twice:

Runs 1 and 6

  1. ClickUp
  2. Teamwork
  3. Xero Projects

Runs 3 and 9

  1. Teamwork
  2. Productive.io
  3. Paymo

The combination of Teamwork, ClickUp and Xero Projects appeared in four runs, but the products were not always presented in the same order.

This suggests that AI Mode may repeatedly return certain candidate groups while still varying the final selection and presentation.

Teamwork’s presence in all ten answers is especially significant. The individual responses changed, but repeated testing revealed that Teamwork had much more consistent visibility than any other product.

The qualification criteria also changed between responses

The variability was not limited to rearranging product names.

The prompt required software that integrated with Xero. In some responses, ClickUp was accepted because it could connect through an automation platform such as Zapier or Make. In another response, ClickUp was excluded because it required paid third-party middleware rather than offering the required native connection.

The meaning of client visibility also appeared to vary. Dedicated client portals, guest accounts, shared project views, progress reports and detailed invoices were sometimes treated as alternative ways to satisfy the same requirement.

These are not necessarily equivalent experiences for the buyer.

This helps explain why the shortlist may change. AI Mode appears to reconstruct the answer each time, and part of that process involves interpreting what qualifies as a suitable solution.

I have not attempted to verify every price, plan feature or integration claim in this experiment. That would require a separate accuracy and citation test.

What can a business learn from one manual search?

One search can answer a limited question:

Did my product appear in this particular response?

It cannot reliably answer:

  • How frequently does my product appear?
  • How often do competitors appear?
  • Which product is most consistently prominent?
  • Does my product disappear from certain versions of the answer?
  • Does its recommendation position change?
  • Does AI Mode describe its features consistently?
  • Is visibility changing over time?

Those questions require repeated observations.

A simple form of recommendation share of voice

This experiment provides a basic manual illustration of recommendation share of voice.

Teamwork occupied 10 of the 30 available recommendation positions. ClickUp occupied seven, while Avaza occupied two.

However, this should not be treated as a permanent or statistically conclusive measurement. The experiment used:

  • one query;
  • one Google account;
  • one location;
  • one AI search platform;
  • ten runs conducted during a single hour.

It shows that variability exists and that repeated testing reveals patterns. It does not prove that every user will receive the same distribution.

A proper monitoring programme would need to repeat multiple commercially relevant prompts over time and potentially across several AI platforms.

Why manual monitoring quickly becomes difficult

Running one prompt ten times was manageable. Expanding the method would create substantially more work.

A business might want to monitor:

  • several different customer scenarios;
  • alternative ways of phrasing each scenario;
  • multiple products or services;
  • several AI platforms;
  • changes from week to week;
  • competitor appearances;
  • citations and linked sources.

Even monitoring 20 prompts across two platforms with five repetitions would produce 200 responses for a single testing period.

Manual testing remains useful for learning how AI recommendations behave and deciding which prompts matter. Once the number of prompts grows, recording and comparing the results consistently becomes much harder.

A useful distinction is:

Manual testing is valuable for investigation. Automated monitoring becomes valuable when repeated measurement is required.

What this experiment demonstrates

The clearest conclusion is simple:

One search showed an answer. Ten searches revealed a pattern.

A product appearing once does not prove that it has consistently strong AI visibility. Equally, failing to appear once does not prove that it is invisible.

In this experiment:

  • six different products appeared;
  • four different products occupied first place;
  • six different shortlist combinations were generated;
  • eight distinct ordered results appeared;
  • only one product was present every time.

Businesses checking how they appear in AI-generated recommendations therefore need to think beyond a one-off search.

The useful question is no longer simply, “Did we appear?”

It is:

“Across the searches that matter to our potential customers, how frequently do we appear, where do we appear, and how does that change over time?”

That is a much more demanding question—but it is also a much more meaningful measure of AI visibility.

Testing conducted on 3 August 2026. Google AI Mode and its results may change over time.

Leave a Comment