The Website Had the Answer — But AI Didn’t Always Find or Interpret It Correctly: A Five-Run Test

A business owner looks at their website with knowledge a customer does not have.

They already know whether the restaurant opens on Bank Holidays.

They know whether children are welcome.

They know whether vegan dishes are available.

They know whether the refurbishment mentioned on the website has finished.

They know how the booking system works.

That creates an easy problem to miss:

A website can look complete to the business because the business unconsciously fills in the gaps from its own knowledge.

A customer cannot do that.

Neither can an AI system with any certainty.

We decided to test this using five Birmingham restaurant websites and one realistic customer situation.

The experiment quickly became more interesting than a simple website audit.

Some restaurants provided extremely clear first-party information.

Others left particular questions unresolved.

But even when the answer was explicitly published on the restaurant’s own website, Google AI Mode did not always retrieve or interpret it consistently.

Across five identical runs:

  • a children’s menu that definitely existed was missed three times before being found;
  • another clearly published children’s menu disappeared for one run after being found correctly three times;
  • a restaurant that explicitly welcomes children was at one point described as poorly suited to them;
  • normal opening hours changed between responses;
  • and some highly specific table-availability claims appeared stronger than the visible evidence supporting them.

The experiment ended up exposing a much wider AI visibility problem:

Business knowledge → Website evidence → AI retrieval → Source selection → AI interpretation → Customer conclusion

A problem can occur at every stage.

The Starting Question: Does the Website Actually Answer the Customer?

Restaurant websites often contain attractive photography, descriptions of the food, information about the chef and broad statements about the dining experience.

But that is not necessarily what a customer needs when making a decision.

Someone might instead need to know:

  • Are you open on this particular day?
  • Are your hours different because it is a Bank Holiday?
  • Can I book online?
  • Is there actually space at the time I need?
  • Can a vegan member of the party eat properly?
  • Are children welcome?
  • Is there a children’s menu?
  • Can I see current prices?
  • Is the information obviously current?

Those questions are much more specific than:

What type of restaurant are you?

So rather than auditing the websites as collections of pages, we tested them against a real customer journey.

This follows the same principle behind our Scenario Suitability Test: test whether the available evidence supports a particular customer’s requirements rather than simply asking whether the business sounds suitable in general.

The Customer Scenario

We fixed the following situation:

A party of six wants dinner in Birmingham at around 7pm on Monday 31 August 2026, the August Bank Holiday. Two of the party are children and one adult is vegan. They want to know whether the restaurant is open, whether they can book, whether everyone can eat there, what the likely cost will be and whether the information appears current.

We selected five restaurants:

  • San Carlo Birmingham
  • The Mayan
  • Pasture Birmingham
  • Gusto Italian Birmingham
  • Browns Birmingham

The restaurants were selected before the detailed audit.

We did not choose them because their websites appeared particularly good or bad.

First We Built a First-Party Baseline

Before asking AI, we needed to know what each restaurant actually said about itself.

The rule was simple:

Do not give the business credit for information the customer cannot establish from its own website.

If normal Monday hours were published but Bank Holiday hours were not, we did not automatically assume they were identical.

If a booking button existed, we distinguished:

The customer can check availability

from:

A table for six at 7pm is definitely available.

If the restaurant owner almost certainly knew children were welcome but the website did not make that clear, we did not fill the gap from common sense.

This produced a useful baseline.

Some sites were particularly good at anticipating practical customer questions.

Pasture

Pasture’s own FAQ directly answered questions about:

  • children’s menus;
  • highchairs;
  • baby-changing facilities;
  • vegan options;
  • dietary requirements;
  • and deposits for larger bookings.

Gusto

Gusto clearly explained:

  • vegan and vegetarian choices;
  • family dining;
  • its children’s menu;
  • allergens;
  • accessibility;
  • and practical booking information.

The Mayan

The Mayan provided an especially useful example because it explicitly separated two ideas:

Children are welcome.

But:

There is currently no dedicated children’s menu.

That is good customer information.

A useful answer does not have to be yes. A clear no can remove uncertainty just as effectively.

Browns

Browns was notable because it published information specifically about Bank Holiday dining rather than simply expecting customers to assume normal Monday arrangements applied.

That matters because the customer’s question was not:

Are you normally open on Monday?

It was:

Are you open on this particular Bank Holiday Monday?

Then We Asked Google AI Mode

We gave Google AI Mode the same customer scenario and the same five restaurants.

The prompt asked it to compare each restaurant on:

  • Bank Holiday opening;
  • availability for six people around 7pm;
  • vegan suitability;
  • child suitability and children’s menus;
  • and visible menu pricing.

We also explicitly told it:

Clearly distinguish between information you can verify and information you are inferring.

And:

Do not assume that normal Monday opening hours automatically apply on a Bank Holiday.

Importantly, we did not restrict AI Mode to the restaurants’ own websites.

It was allowed to use whatever online sources it considered useful.

That more closely reflects a real AI search, where the answer may combine:

  • first-party websites;
  • booking platforms;
  • directories;
  • social media;
  • restaurant-festival sites;
  • reviews;
  • and other sources.

We then repeated the identical prompt five times.

As with our other manual experiments, five runs are not a statistically significant study. They are a practical exploratory check designed to reveal whether obvious variation appears when the same query is repeated.

Our broader approach is explained in How Many Times Should You Repeat an AI Search When Testing Brand Visibility?.

The Five Runs Produced a Very Different Picture

Several of the most useful facts can be summarised simply.

Fact TestedRun 1Run 2Run 3Run 4Run 5First-Party Position
San Carlo has a kids’ menuNoNoNoYesYesYes
Pasture has a kids’ menuYesYesYesNoYesYes
The Mayan welcomes childrenUnclearUnclearPoor fitYesYesYes
The Mayan has a kids’ menuNo foundNo foundNoNoNoExplicitly no
Gusto has a kids’ menuYesYesYesYesYesYes

The websites did not need to change for these different conclusions to appear.

We made no changes to them, and during the short testing period we found no evidence that the relevant first-party information itself changed.

The AI responses did.

San Carlo: The Answer Was There, But AI Missed It Three Times

San Carlo has a dedicated children’s menu.

Its own website explicitly publishes its Menu Per Bambini and identifies Birmingham among the locations where it is available.

Yet the first three AI Mode runs failed to recognise it.

Only Runs 4 and 5 found the menu correctly.

That produces an important distinction:

Publishing explicit first-party information gives AI evidence it can use. It does not guarantee that AI will retrieve or use that evidence every time.

If we had stopped after one search, we might have concluded that San Carlo’s website had failed to explain its children’s offering.

It had not.

The information existed.

The AI had failed to use it.

San Carlo also produced a simpler factual fluctuation.

In one response its normal Monday closing time was stated incorrectly.

A later response corrected it.

So the variability extended beyond recommendations or interpretation into ordinary business facts.

Pasture: Clear Information Can Disappear

Pasture produced almost the reverse pattern.

Its website is unusually explicit about children.

Its own FAQ clearly says it has:

  • a children’s menu;
  • highchairs;
  • baby-changing facilities;
  • and vegan options.

AI correctly recognised the children’s menu in Runs 1, 2 and 3.

Then Run 4 suddenly said:

Not Recommended / No Kids’ Menu

and claimed there was no evidence of a dedicated children’s menu online.

Run 5 returned to the correct answer.

The pattern was therefore:

Yes → Yes → Yes → No → Yes

That is particularly useful because it prevents a simple explanation such as:

San Carlo’s information was just difficult for AI to find.

Pasture’s information was extremely explicit and had already been successfully retrieved three times.

Then it disappeared from the fourth response.

The Mayan: The Facts Were Right, But the Interpretation Wasn’t

The Mayan gave us one of the strongest examples of an interpretation problem.

Its own FAQ clearly separates three facts:

Children are welcome.

Highchairs are available.

There is no dedicated children’s menu.

That should lead to a fairly straightforward conclusion:

Suitable for children, but no separate kids’ menu.

Instead, one AI response described child suitability as:

Poor / Inferred No

because no children’s menu was available and the restaurant was portrayed as having a more adult-focused atmosphere.

That changes the meaning.

No children’s menu does not mean children are not welcome.

Later runs eventually produced the more accurate interpretation.

This is closely related to what we call an Interpretation Gap.

The relevant information exists.

The weakness lies in the conclusion drawn from it.

For a real customer, that difference matters.

A parent seeing:

Poor suitability for children

might dismiss the restaurant.

Yet the restaurant itself explicitly welcomes children.

The Mayan Also Produced a Straightforward Factual Error

Later AI runs gave The Mayan’s Monday hours as 12pm to midnight.

Its own current first-party information states normal Monday opening as 5pm to 11pm.

This is useful because it separates two different types of problem.

Retrieval or factual failure

The first-party information exists, but the AI gives the wrong fact.

Interpretation failure

The facts are substantially correct, but the conclusion drawn from them is wrong.

The Mayan produced examples of both.

Gusto: The Stable Control Case

Not everything fluctuated.

Gusto’s treatment was remarkably consistent.

Across all five runs, AI correctly understood that it:

  • offered vegan choices;
  • catered for families;
  • had a dedicated children’s menu;
  • and published useful pricing information.

That matters.

Our conclusion is not:

AI cannot understand restaurant websites.

Clearly it can.

The more interesting finding is that:

some facts were represented consistently while others changed substantially between identical runs.

Why that happens is a separate question.

Our experiment does not establish the mechanism.

But it shows why one-off testing can miss the instability.

Browns: Specific Information Reduced the Need for Guesswork

Browns provided another useful contrast.

Instead of publishing only normal Monday hours, it also had content specifically addressing Bank Holiday dining.

That made it much easier to support the broad conclusion that the restaurant intended to trade over the Bank Holiday weekend.

This illustrates an important principle.

Compare:

Monday: 9am–11pm

with:

We are serving customers over the Bank Holiday weekend.

The second statement answers the customer’s actual situation much more directly.

It reduces the amount of inference required.

That does not necessarily prove that every normal Monday hour applies unchanged.

But it gives both the customer and AI stronger evidence about the particular circumstance being asked about.

This connects with our wider testing around customer-need themes.

Businesses often provide broad category information.

Customers frequently need something much more specific.

Live Availability Introduced Another Type of Confidence Problem

Several AI responses became increasingly precise about table availability.

They claimed particular slots such as:

  • 18:45;
  • 19:00;
  • 19:15;
  • and other neighbouring times.

In one response, one specific time was even described as fully booked while nearby times remained available.

Those are very precise claims.

But there is an important difference between:

This restaurant has an online booking system

and:

I have verified that a table for six is available at 7pm on 31 August 2026.

The second requires evidence that the actual reservation system has been queried using:

  • the correct restaurant;
  • the correct date;
  • six guests;
  • and the relevant time.

In at least one response, the visible citation trail did not clearly support all of the exact availability claims attached to it.

So for this experiment we distinguish:

AI claimed availability was confirmed

from:

we independently verified that exact table availability ourselves.

We did not complete bookings, so we should not treat those as equivalent.

This connects directly with another lesson from our citation testing:

A citation beside an AI answer does not automatically prove that the source supports every nearby claim.

We explore that separately in Being Cited by AI Does Not Mean It Used Your Page Accurately.

AI Did Not Always Use the Strongest Available Source

Another pattern appeared repeatedly.

The restaurants often provided the answer themselves.

Yet AI Mode also used information from sources such as:

  • OpenTable;
  • TripAdvisor;
  • TheFork;
  • Instagram;
  • Facebook;
  • festival websites;
  • directories;
  • and other third-party pages.

Third-party sources are not inherently a problem.

Sometimes they may provide useful additional evidence.

But there were cases where AI appeared to rely on indirect or weaker evidence despite clearer first-party information being available.

That suggests AI visibility testing should ask more than:

Was the final answer correct?

We may also need to ask:

What evidence was used to reach it?

There are at least three useful categories.

Correct answer, strong evidence

The ideal outcome.

Correct answer, weak or indirect evidence

The customer receives approximately the right conclusion, but the evidential pathway is less robust.

Incorrect answer despite stronger evidence being available

The business has supplied the information, but AI fails to retrieve or interpret it correctly.

Our five runs produced examples of all three.

What the Five Runs Showed

The experiment ultimately revealed several different types of behaviour:

OutcomeExample
Stable correct interpretationGusto’s children’s information
Repeated missed first-party evidenceSan Carlo’s kids’ menu in Runs 1–3
Previously found evidence disappearedPasture’s kids’ menu in Run 4
Correct answer reached without using the clearest evidenceEarlier Mayan children’s-menu responses
Correct facts interpreted incorrectlyNo Mayan kids’ menu became poor child suitability
Straight factual errorThe Mayan’s Monday hours
Fact corrected between identical runsSan Carlo’s normal hours
Confidence changed between runsBank Holiday conclusions moved between inferred and verified
Specific claim stronger than clearly visible evidenceSome live availability claims

One AI response would have exposed only a fraction of this.

Repeated testing showed that the failure itself could move around.

Limitations of the Experiment

There are several reasons to be careful about what this test proves.

Five runs are exploratory

Five repeated searches are useful for revealing obvious variation.

They do not establish statistically representative error rates.

Five restaurants do not represent the whole restaurant industry

This was a practical experiment using five Birmingham restaurants.

We should not infer that the same proportion of issues exists across restaurant websites generally.

AI was allowed to use third-party sources

Our baseline focused on what the restaurants themselves published.

AI Mode was allowed to search more widely.

That was deliberate, but it means the experiment tested the whole online evidence environment rather than first-party websites alone.

Live availability was not independently booked

We recorded what AI claimed.

We did not complete reservations ourselves.

Sequential testing may matter

Some later responses appeared to improve.

San Carlo’s children’s menu went:

missed → missed → missed → found → found

The Mayan’s child-suitability interpretation also became better in later runs.

Could repeated searches have influenced later answers?

We cannot completely rule out effects from session history, personalisation or other context.

But the results do not look like simple cumulative learning.

Pasture went:

correct → correct → correct → wrong → correct

If the system were simply learning the correct facts from each previous run, we would not expect an already-correct answer to disappear in Run 4.

The safer conclusion is:

Retrieval, source selection and interpretation changed across repeated searches, but this experiment cannot determine how much of that variation came from stochastic behaviour, changing retrieval, search history, personalisation or other system effects.

The experiment does not prove that more detailed websites improve AI rankings

That was not what we tested.

Our result is narrower.

Clear first-party information gives AI something useful to work with.

It does not guarantee selection, citation or correct interpretation.

What This Means for Businesses

The original website lesson remains important.

Businesses often review websites from the perspective of someone who already knows the business.

A more useful test is:

Could someone who knows nothing about us answer this customer’s question using only what we actually publish?

For a restaurant, that might include:

  • ordinary opening hours;
  • exceptional Bank Holiday hours;
  • current menus;
  • dietary options;
  • children’s facilities;
  • group booking requirements;
  • accessibility;
  • deposits;
  • refurbishment closures;
  • reopening dates;
  • and live booking pathways.

For an accountancy firm, the equivalent information might include:

  • industries served;
  • accounting software supported;
  • payroll;
  • management accounts;
  • fixed-fee arrangements;
  • business size;
  • locations covered;
  • specialist circumstances;
  • and how a prospective client actually engages the firm.

For a dentist, it might include:

  • nervous-patient support;
  • sedation;
  • treatment options;
  • pricing;
  • accessibility;
  • and whether new patients are being accepted.

This is not really an AI optimisation trick.

It is better first-party information.

But Publishing the Information Is Only the First Step

This experiment added a second lesson.

We can think about the journey like this:

Business knowledge → Website evidence → AI retrieval → Source selection → AI interpretation → Customer conclusion

Each stage can fail.

Business knowledge → Website evidence

The business knows something important but never publishes it.

Website evidence → AI retrieval

The answer exists, but AI does not find or use it.

San Carlo demonstrated this repeatedly.

AI retrieval → Source selection

Clear first-party evidence exists, but AI uses weaker or more indirect sources.

Source selection → AI interpretation

The facts are substantially correct, but AI draws an overly broad conclusion.

The Mayan provided a clear example.

AI interpretation → Customer conclusion

The final response becomes more confident or specific than the evidence clearly justifies.

The live table-availability claims highlighted this risk.

That makes AI visibility testing much broader than simply counting mentions.

The useful questions become:

Does AI find the right evidence about the business?

Does it use the strongest available source?

Does it preserve important distinctions?

Does it describe the business accurately?

And does it do those things consistently when the same customer situation is tested again?

The Bigger AI Visibility Lesson

We began this experiment with a simple question:

Does the business website actually answer a real customer’s questions?

We ended with a more complicated one:

If the website contains the answer, can AI reliably carry that information all the way through to the customer?

Across five exploratory runs, the answer was:

sometimes yes, sometimes no.

The important result was not simply that AI made mistakes.

It was that the mistakes themselves changed while the relevant first-party information remained available.

A children’s menu could exist throughout the experiment and be missed repeatedly.

Another could be found three times, disappear once and return.

A restaurant could explicitly welcome children while the AI interpreted the lack of a children’s menu as poor child suitability.

Basic opening hours could change between answers.

And highly specific live-availability claims could be presented with more certainty than the visible supporting evidence appeared to justify.

For businesses, the first step remains clear:

Make customer-important information explicit.

Do not expect customers—or AI—to fill gaps using knowledge that only the business possesses.

But then test what happens after publication.

Because this experiment suggests that even when the answer is sitting on the website, AI may not always find it, may not always use the strongest evidence, and may not always interpret it in the same way.

Leave a Comment