Checking whether a business appears in an AI-generated response sounds straightforward:
- Enter a relevant query.
- Read the answer.
- Record whether the business appeared.
The problem is that one response may not represent the wider pattern.
In my ten-run Google AI Mode experiment, I submitted exactly the same product-selection query ten times. Six different products appeared, four occupied first place, and only one was recommended in every response.
A useful manual visibility test therefore requires repeated searches and a consistent method. This guide explains a practical process for establishing an initial baseline and monitoring how the results change over time.
Decide what you are measuring
Begin by defining the purpose of the test.
You might want to know:
- whether a business or product is mentioned;
- whether it is recommended as a solution;
- where it appears in an ordered shortlist;
- which competitors are included;
- whether the business’s website is cited;
- how the product is described;
- how frequently it appears across repeated responses.
These questions require different levels of recording.
A product-recommendation test may only require names, positions and competitors. A citation test would also need the linked pages and the claims associated with them.
Deciding this in advance prevents a simple experiment from becoming an unnecessarily large investigation.
Select a commercially meaningful prompt
A repeatable method cannot rescue an irrelevant prompt.
Choose a search representing a genuine situation in which a potential customer might discover, compare or select what the business offers.
My earlier experiment found that adding time tracking, client access, Xero integration and a £100 budget produced a very different software shortlist from a broad project-management query. The full method is explained in How to Choose AI Search Prompts That Actually Matter to Your Business.
A useful question to ask is:
If my business appeared in this response, would the person asking the question be a plausible customer?
If the answer is no, the prompt may have limited commercial importance even if it generates a mention or citation.
Lock the prompt and testing conditions
Save the exact prompt before beginning.
Keep its wording, punctuation, customer details, constraints and requested response format unchanged throughout the test. Copying and pasting the saved version is safer than retyping it.
If you later improve the wording, treat the revised prompt as a new test. Changing it halfway through makes the earlier and later responses less directly comparable.
You should also record the testing conditions:
- AI platform;
- date and approximate time;
- country or location;
- whether you were signed in;
- whether personalisation was enabled;
- model or mode selected, if a choice was available;
- browser or device, if relevant.
No AI search is completely neutral. Results may be influenced by location, account settings, previous activity and changes within the platform.
The practical objective is to keep conditions as consistent as reasonably possible and document them honestly.
Is incognito mode necessary?
Incognito mode can reduce the influence of some saved browser activity, but it does not remove every contextual signal or create a perfectly neutral search.
Some AI services may also behave differently when the user is signed out.
A reasonable manual baseline might use:
- the same account and browser;
- personalisation turned off where possible;
- the same location;
- fresh searches;
- documented testing conditions.
Another business might deliberately test the normal personalised experience. That is also valid if it matches the objective and is applied consistently.
Avoid switching unpredictably between signed-in, signed-out, personalised and incognito searches while treating all the results as equivalent.
Decide the number of repetitions before starting
There is no universal correct number of runs.
| Number of runs | Possible purpose |
|---|---|
| 1 | A spot check, not a measurement |
| 3 | A small exploratory test |
| 5 | A more useful initial indication |
| 10 | A clearer view of recurring patterns |
| Scheduled repeated tests | Monitoring longer-term change |
Choose the number before seeing the results. Stopping only when a preferred business appears would weaken the test.
One run establishes what happened in that particular response. Several runs make it possible to observe appearance frequency, changing positions and recurring competitors.
Run each search independently
If you are testing repeated independent responses, begin every run as a fresh search rather than continuing the previous conversation.
Follow-up questions may be influenced by the existing response and conversational context.
Use the following process:
- Open a fresh search.
- Paste the exact saved prompt.
- Select the required AI mode or platform.
- Allow the response to finish.
- Do not ask a follow-up question.
- Record the result and save a screenshot.
- Start the next run separately.
- Continue until the predetermined number of runs is complete.
Capture each response before clicking its source links or conducting related searches.
The objective is not to make the platform forget everything about the user. It is to avoid deliberately carrying one generated answer into the next test.
What should you record?
A simple spreadsheet can contain one row for every run.
| Field | Information to save |
|---|---|
| Prompt ID | A reference such as PM-01 |
| Exact prompt | Complete unchanged wording |
| Platform | Google AI Mode, ChatGPT or another platform |
| Date | Testing date |
| Account condition | Signed in or signed out |
| Personalisation | On or off |
| Run number | 1, 2, 3 and so on |
| Brand present | Yes or no |
| Recommendation position | First, second, third or not shown |
| Competitors | Other businesses or products shown |
| Screenshot | Saved filename |
| Notes | Any unusual result |
Screenshots preserve evidence of what appeared at the time. A structured table makes the overall pattern easier to understand.
For example:
| Run | First | Second | Third |
|---|---|---|---|
| 1 | Product A | Product B | Product C |
| 2 | Product B | Product A | Product D |
| 3 | Product A | Product D | Product B |
You do not need to publish every complete response. Retain the full evidence, then present the information relevant to the question being tested.
Do you need to record every citation?
Only map citations if citation visibility forms part of the experiment.
A recommendation test might require only:
- products shown;
- recommendation order;
- appearance frequency;
- competitors;
- screenshots.
A citation test may additionally require:
- citation marker;
- linked page and domain;
- claim associated with the citation;
- whether the page supports that claim.
Mapping every source can become cumbersome, so do not turn every visibility check into a full citation audit without a clear reason.
Repetition establishes a baseline
Several identical searches conducted close together show how variable an AI response is under broadly consistent conditions.
My ten Google AI Mode runs took approximately one hour. The wider web was unlikely to have changed substantially during that period, but the recommendations still varied.
Those immediate repetitions created an initial baseline.
Suppose a product appears in eight of ten responses and occupies first position four times. That result can later be compared with another controlled test using the same prompt and methodology.
One isolated search cannot provide that context.
Monitoring tracks change over time
The online information environment does not remain fixed.
Over time:
- new websites and products appear;
- existing pages are updated or removed;
- prices and product features change;
- businesses publish new information;
- search engines discover and index new pages;
- competitors change their positioning;
- AI models and retrieval systems are updated;
- algorithms and interfaces evolve.
A prompt may produce a different pattern next month because the underlying landscape has changed, not merely because an individual response varied.
This creates an important distinction:
Repetition shows how variable an AI answer is today. Monitoring shows how visibility changes as the web and AI systems evolve.
After establishing a baseline, repeat the same controlled test at a sensible interval.
Possible starting points include:
- weekly for important prompts in fast-changing markets;
- monthly for core commercial prompts;
- quarterly for smaller businesses conducting manual reviews;
- after a major product launch, price change or website restructure.
The right frequency depends on the market, commercial importance of the prompts and available resources. The objective is to identify sustained patterns rather than react to every individual movement.
When does automation become useful?
Manual testing is valuable for learning:
- which prompts matter;
- how variable responses are;
- what information needs recording;
- which competitors repeatedly appear;
- which measurements are commercially useful.
A small business checking a few prompts may not require specialist monitoring software.
However, the workload grows rapidly. Monitoring 20 prompts across two AI platforms with five repetitions would produce 200 responses during each testing period.
That is before checking citations, comparing descriptions or maintaining historical reports.
Automation becomes more useful when a business needs to track:
- numerous prompts;
- several AI platforms;
- multiple products or locations;
- recommendation positions;
- competitors;
- citations;
- historical trends;
- regular client or management reports.
The manual process helps establish what deserves to be monitored before that process is scaled.
What can the test prove?
A controlled manual test can reveal:
- recurring appearances;
- recommendation variability;
- prominent competitors;
- differences between prompts;
- changes between testing periods.
It cannot prove:
- what every individual user will see;
- why the AI selected a particular business;
- that an appearance rate will remain permanent;
- that every generated claim or citation is accurate;
- that one particular website change caused a movement.
AI-generated answers are dynamic. A repeatable method does not remove that uncertainty. It makes the uncertainty easier to observe and compare.
The practical conclusion
A single search is an observation. Several controlled runs create an initial pattern. Repeating the same test later turns that pattern into monitoring.
The method does not need to be technically complicated. It needs commercially meaningful prompts, consistent conditions and an honest record of what happened.
That allows a business to move beyond asking:
“Did we appear when I searched today?”
The more useful question is:
“How consistently do we appear for the customer situations that matter, and is that visibility changing over time?”