A business can update the information on its own website and still have older or conflicting information elsewhere online.
Opening hours are a simple example.
A restaurant might show one set of hours on its own website, while Google, OpenTable, TripAdvisor or a local directory shows something different.
That raises an important AI visibility question:
When first-party and third-party sources disagree, which version does AI use?
We tested this with three Birmingham restaurants where we had already identified genuine opening-hours conflicts.
We then asked Google AI Mode the same narrow factual question five times for each restaurant.
That gave us 15 responses in total.
The result was not that AI always trusted the official website.
It was not that AI always trusted third-party platforms either.
Instead, the three restaurants produced three distinctly different patterns.
| Restaurant | Runs Matching Current First-Party Information | Conflict Detected | Overall Pattern |
|---|---|---|---|
| San Carlo Birmingham | 5/5 | 5/5 | Stable answer and stable conflict awareness |
| Pasture Birmingham | 5/5 | 1/5 | Stable answer, but conflict usually went unnoticed |
| The Mayan Birmingham | 2/5 | 3/5 | Answer and conflict handling both varied |
That difference is the real story of the experiment.
We Identified the Conflicts Before Testing AI
This is important.
We did not run repeated AI searches, notice errors and then go looking for conflicting sources afterwards.
We first established the published evidence.
These were naturally occurring disagreements. We did not manufacture or alter any information.
San Carlo Birmingham
San Carlo’s own website showed Monday hours of:
12:00–22:00
But third-party sources showed alternatives including:
12:00–21:30
and:
12:00–23:00
The Mayan Birmingham
The Mayan’s own Contact and FAQ pages showed Monday opening as:
17:00–23:00
But other online sources showed:
12:00–00:00
That is a substantial difference for someone considering a Monday lunch.
Pasture Birmingham
Pasture’s own website showed:
12:00–15:00 and 17:00–21:30
So there is a break between lunch and dinner.
Other sources presented different versions, including:
12:00–21:30 continuously
and:
12:00–15:00 only
This gave us three very clean examples where published sources disagreed over a basic factual question.
The Prompt
For each restaurant we used the same format.
For example:
What are the current opening hours for San Carlo Birmingham on a Monday? Please verify the information online and cite the sources you use. If there is any uncertainty about the hours, make that clear.
Only the restaurant name changed.
Each run was started in a fresh AI Mode conversation to reduce direct conversational carryover.
That does not eliminate every possible effect from search history, personalisation or other system behaviour, so these should still be treated as practical exploratory observations rather than fully independent laboratory trials.
The aim was much simpler:
Would AI handle the same known source conflict consistently when the same factual question was repeated?
San Carlo: Stable Across All Five Runs
San Carlo produced the cleanest result.
Across all five runs, AI Mode gave:
12:00–22:00
That matched San Carlo’s current first-party information every time.
AI also noticed the conflicting third-party schedules in every run.
| Run | Matches First-Party Information? | Conflict Detected? |
|---|---|---|
| 1 | Yes | Yes |
| 2 | Yes | Yes |
| 3 | Yes | Yes |
| 4 | Yes | Yes |
| 5 | Yes | Yes |
So for this particular source conflict, AI handled the situation very consistently.
It repeatedly treated the restaurant’s own website as the strongest source.
But the reasoning around the disagreement was less stable.
In some runs, AI suggested that the earlier third-party closing time might represent:
- last orders;
- final reservations;
- or an earlier kitchen closure.
Those explanations were plausible.
But the evidence we had did not establish that they were true.
Other runs handled the same disagreement more cautiously and simply advised checking directly if visiting close to closing time.
That gives us our first useful distinction:
Answer stability and reasoning stability are not necessarily the same thing.
Pasture: Stable Answer, Unstable Conflict Awareness
Pasture also matched the current first-party information in all five runs.
Every response gave:
12:00–15:00 and 17:00–21:30
On simple factual matching, that looks perfect.
But something else was happening.
AI noticed the known conflicting information in only one of the five runs.
| Run | Matches First-Party Information? | Conflict Detected? |
|---|---|---|
| 1 | Yes | No |
| 2 | Yes | No |
| 3 | Yes | Yes, partially |
| 4 | Yes | No |
| 5 | Yes | No |
Several responses effectively presented the split hours as settled information with no meaningful uncertainty.
Yet we already knew that other prominent sources showed different schedules.
That leads to another important distinction:
A final answer can match the first-party source without AI having recognised or resolved the wider evidence conflict.
For a customer, the answer may still be useful.
For a business trying to understand how it is represented across AI search, the picture is less complete.
The conflicting information still exists online.
AI just did not surface it in most runs.
The Mayan: Correct, Wrong and Unresolved
The Mayan produced the strongest example in the experiment.
Its current first-party Monday hours were:
17:00–23:00
But the five identical AI Mode searches produced three different treatments of the same source conflict.
| Run | AI Answer | Matches First-Party Information? | Conflict Handling |
|---|---|---|---|
| 1 | 17:00–23:00 | Yes | Conflict detected, official site preferred |
| 2 | 17:00–23:00 | Yes | Conflict detected, official site preferred |
| 3 | 12:00–00:00 | No | Conflict missed, high confidence |
| 4 | 12:00 or 17:00 unclear | — | Conflict detected, left unresolved |
| 5 | 12:00–00:00 | No | Conflict missed, presented as verified |
The sequence was:
First-party → First-party → Conflicting version → Unresolved → Conflicting version
The prompt had not changed.
The underlying source disagreement had not been introduced by us.
Yet the answer did.
The Wrong Answers Were Not Necessarily Less Confident
This was one of the most striking parts of the test.
In Runs 1 and 2, AI recognised that the sources disagreed and used cautious language.
Run 4 also recognised the conflict and admitted uncertainty.
But Runs 3 and 5 failed to recognise the contradiction and became more confident.
Run 3 described the midday-to-midnight schedule with high certainty.
Run 5 went further and presented the information as:
Source Verified
That gives us an important observation:
In this experiment, confidence was not a reliable proxy for evidential strength.
The more confident answer was not necessarily the better-supported one.
The Mayan Run 5: Why Checking Citations Matters
Run 5 may be the clearest example in the entire test.
AI stated that The Mayan was open on Monday from:
12:00 PM to 12:00 AM
It then said that this had been confirmed via The Mayan’s own website.
But the current first-party FAQ we had recorded before testing showed:
17:00–23:00
So this was not simply:
AI chose a third-party source over the restaurant.
It was:
AI presented the restaurant’s own website as verification for a statement that did not match the information on that page.
That matters because citations create confidence.
A user may reasonably assume that an official-source citation means the claim has been checked against the business’s own information.
This example shows why that assumption can be unsafe.
It connects directly with our earlier test:
Being Cited by AI Does Not Mean It Used Your Page Accurately
The practical lesson is simple:
If the information matters, open the citation and check what the source actually says.
AI Also Tried to Explain Why Sources Disagreed
Another recurring behaviour appeared across the experiment.
When AI recognised conflicting sources, it sometimes tried to reconcile them.
With San Carlo, it suggested that an earlier closing time might represent last orders or final reservations.
With The Mayan, different runs suggested different possible explanations involving lunch service and website updates.
These explanations often sounded reasonable.
But reasonable is not the same as established.
There is an important difference between:
These sources disagree.
and:
This is why they disagree.
The first may be directly supported by the evidence.
The second may be an inference.
When AI quietly turns that inference into a confident explanation, a new layer of uncertainty is introduced.
The 15-Run Result
Across the 15 searches:
- 12 responses matched the current first-party information;
- 2 gave the conflicting Mayan version;
- 1 left The Mayan conflict unresolved.
But that headline result hides much of what mattered.
San Carlo consistently recognised and resolved the conflict.
Pasture consistently matched the first-party information while usually failing to recognise the conflict at all.
The Mayan moved between first-party information, conflicting information and unresolved uncertainty.
So the experiment did not show that AI always favours first-party information.
It also did not show that AI is incapable of resolving conflicting information.
Instead:
The handling of the conflict itself varied between restaurants and, in The Mayan’s case, between identical runs.
Four Things Worth Measuring Separately
The test suggests that conflicting-information audits need more than a simple right-or-wrong score.
1. Answer Accuracy
Does the AI answer match the current first-party information?
This is the obvious first check.
2. Conflict Awareness
Does AI recognise that other credible sources say something different?
Pasture shows why this matters.
It matched the first-party information five times, but noticed the known conflict only once.
3. Source Fidelity
Does the cited source actually support the claim AI attaches to it?
The Mayan Run 5 demonstrates why this should be checked separately.
A citation being present is not enough.
4. Conflict Handling
When sources disagree, does AI preserve uncertainty appropriately?
Does it choose the better-supported source?
Does it explain the disagreement cautiously?
Or does it construct a plausible but unverified story about why the sources differ?
These are different qualities.
A response can perform well on one and badly on another.
Why AI Warnings Matter
AI-generated search responses commonly carry warnings that mistakes can occur and that important information should be checked.
It is easy to read that as a generic disclaimer.
The Mayan example shows why it has practical meaning.
A customer could ask exactly the same question and receive:
5pm opening
on one occasion, and:
12pm opening
on another.
The second answer could even be presented with high confidence and an official citation.
If the user actually opened the cited page, the contradiction would become apparent.
So for someone using AI:
Check important information independently, or at least check the cited source yourself.
That is especially sensible for information where acting on the wrong answer could matter.
AI search is still developing and is likely to improve considerably over time.
This experiment should therefore not be read as evidence that AI Mode cannot handle conflicting information.
San Carlo shows that it sometimes handles it extremely well.
The narrower lesson is that mistakes, source conflicts and inconsistent interpretations can still occur.
The Business Lesson Is Different
For businesses, there is another implication.
Do not assess your AI representation from a single successful search.
If we had tested The Mayan once, we would have concluded that AI handled the source disagreement perfectly.
The same would have been true after Run 2.
Only repeated testing exposed:
correct → correct → conflicting → unresolved → conflicting
That reinforces one of the core principles behind our testing:
One AI answer is an observation, not a measurement.
Businesses should therefore think about repetition in two ways.
Repeat Important Prompts
Ask the same commercially important question several times.
That reveals whether the representation is stable.
Test Several Prompts Around the Same Customer Need
Customers will not all phrase their questions identically.
Someone might ask:
What time does this restaurant open on Monday?
Another might ask:
Which Birmingham restaurants are open for lunch on Monday?
Another might ask:
Where can I eat near The Mailbox at 1pm on Monday?
Those different prompts may cause different sources to become important.
That is why we recommend building tests around real customer situations rather than trying to test an unlimited list of arbitrary search phrases.
A business does not need to run hundreds of manual prompts.
A practical starting point is:
Choose a small number of commercially important customer situations, create several realistic prompts around them, and repeat the most important ones enough times to see whether the answers remain stable.
Businesses Should Also Check the Wider Evidence Environment
Keeping the business website accurate remains essential.
But this experiment suggests it may also be sensible to check prominent third-party sources.
Depending on the business, that could include:
- Google Business Profile;
- booking platforms;
- major directories;
- industry listings;
- travel or review sites;
- and important social profiles.
Updating your own website does not necessarily remove older information elsewhere.
AI may encounter both versions.
Sometimes it will resolve the difference well.
Sometimes it may not.
What This Experiment Shows
We began with:
When a business website and third-party listings disagree, which does AI believe?
The answer after 15 runs is:
It depends — and the way AI handles the disagreement can itself vary.
Sometimes AI consistently favours first-party evidence.
Sometimes it gives the first-party answer without noticing the conflict.
Sometimes it leaves the conflict unresolved.
Sometimes the conflicting third-party version becomes the answer.
And sometimes an official citation can appear to support a claim that the official source does not actually contain.
That means AI visibility testing should not stop at:
Did AI mention us?
or even:
Did AI get us right once?
A stronger test asks:
Does the answer match our current information?
Does AI recognise conflicting sources?
Do its citations really support what it says?
And does the representation remain stable when the same important customer question is repeated?
For users, the lesson is to verify important information.
For businesses, the lesson is slightly different:
Test how consistently AI represents you — not simply whether it gets you right once.