We Built a Page for an AI Search Query — Then Tested What It Took for AI Mode to Cite It

We published a page specifically designed to answer a question we had already tested in Google AI Mode.

Google Search Console confirmed that the page was indexed within roughly ten minutes of our indexing request.

We then repeated the original query ten times.

AI Mode did not cite the page once.

We made the user’s information need progressively more specific. The answers changed substantially, but our page still did not appear.

Finally, we asked AI Mode to find a resource addressing the particular distinction our page had been created to explain.

This time it found and cited the exact page.

That progression may tell us considerably more about AI visibility than the citation itself.

What We Wanted to Test

A simple theory of AI visibility might look something like this:

Create a page that answers a question well → get it indexed → become a possible source for AI-generated answers.

But being eligible to appear and actually being selected are clearly not the same thing.

We wanted to see what happened when we controlled the sequence as much as was practical:

  1. Establish what AI Mode said before the resource existed.
  2. Create a page specifically addressing the question.
  3. Publish it and request indexing.
  4. Confirm that Google had indexed it.
  5. Repeat the unchanged query.
  6. Gradually make the information need more specific.
  7. See whether AI Mode ever selected the new page.

This was not an attempt to reverse-engineer Google’s ranking system.

It was a practical observation test.

Test 1: Start With a Broad Question

Our original query was:

If I want to test whether my business is consistently recommended by Google AI Mode, how many times should I run the same query and what should I record?

Before creating the new resource, we ran that identical query five times.

The headline answer was remarkably stable.

Baseline RunRecommended Number of Searches
130
230
330
430
530

Five runs produced five recommendations of exactly 30 searches.

AI Mode repeatedly described 30 as something close to a statistical minimum, using explanations involving statistical reliability, significance or establishing a dependable baseline.

The number was stable.

The surrounding methodology was not.

Different runs suggested completing the tests over anything from a few days to around two weeks. They also disagreed about whether locations, devices, accounts and browser states should be held constant or deliberately varied.

So we had already found an interesting problem:

AI Mode appeared highly confident about 30, but considerably less settled about what those 30 observations were supposed to represent.

We Created a Resource to Address That Problem

We then wrote:

How Many Times Should You Repeat an AI Search When Testing Brand Visibility?

The page did not simply replace 30 with another supposedly correct number.

Its central argument was that the appropriate testing method depends on what you are trying to learn.

It separated three different objectives:

  • a quick manual repeatability check;
  • wider AI visibility monitoring across multiple customer queries;
  • formal statistical estimation.

For a simple repeatability check, we suggested five identical runs as a practical starting point.

But we explicitly said that five runs are not statistically significant.

The purpose is much narrower:

Can we see obvious fluctuation when the same query is repeated under broadly similar conditions?

That is different from attempting to estimate the true probability of a business appearing with a defined level of statistical confidence.

The page therefore challenged an important part of the existing 30-run answer while still recognising that larger samples may be appropriate for different kinds of study.

Google Confirmed the Page Was Indexed Within Minutes

After publication, we requested indexing through Google Search Console.

When we checked roughly ten minutes later, Search Console displayed:

URL is on Google

and:

Page is indexed

We cannot identify the exact second Google added the page to its index.

What we can say is that Search Console had confirmed indexing within roughly ten minutes of our request.

That gave us a particularly clean sequence:

Before the page existed: five baseline tests.

Intervention: publish the new resource.

Indexing: confirmed.

After indexing: repeat exactly the same query.

Ten Broad-Query Tests After Indexing

We then ran the unchanged query another ten times.

Post-Index RunRecommended Number of Searches
130–50
2At least 30
330
415
510
6At least 30
7At least 30
810–15
9At least 30
10At least 30

This revealed something useful even before considering our new page.

The original:

30 / 30 / 30 / 30 / 30

had looked like an extremely robust consensus.

The next ten answers ranged from 10 to 50.

Thirty remained the clear centre of gravity, but it was no longer an invariant answer.

That is another reason repeated AI testing matters. Even an apparently perfect five-run consensus can prove less durable when testing continues.

Our guide to running a repeatable manual AI visibility test is built around exactly this principle: one AI response should not automatically be treated as representative.

But the main result for this experiment was simpler.

Across those ten post-indexing runs:

AI Visibility Testing mentions: 0/10

New resource citations: 0/10

The page was indexed.

It directly addressed the question.

But we found no evidence that AI Mode was selecting it for the broad query.

Test 1 Result: Indexed and Relevant Was Not Enough

This is an important distinction.

Indexing made the page available to Google.

Relevance meant the page clearly addressed the subject.

Neither guaranteed selection.

That led us to start thinking in terms of what we called a candidate pool.

This is our own working model, not Google’s terminology, and we cannot see Google’s internal retrieval process.

The idea is simply useful for thinking about the experiment.

For any information need, there may be many pages that could plausibly contribute to an answer.

A page may therefore have to do more than qualify as relevant.

It must compete with all the other plausible sources.

And that competition is likely to depend on more than the number of pages.

It could also depend on:

  • the quality of those competing sources;
  • how closely each source matches the particular requirement;
  • the amount of existing agreement around an answer;
  • the evidence supporting different claims;
  • how much genuinely distinctive information each source contributes.

Our broad AI-visibility question potentially had an enormous competitive environment.

There are many pages discussing AI visibility, AI monitoring, Share of Voice, sampling, search testing, AI optimisation and statistical analysis.

Our new page was one relevant resource among many.

Compare That With Our Nervous-Patient Dentist Experiment

This contrasted sharply with an earlier experiment:

Would AI Mode Find the Dental Practice With the Strongest Evidence? A Five-Run Test

That customer was not merely searching for a dentist.

They needed a private dentist in Leeds who could meet several connected requirements involving severe anxiety, previous bad experiences, non-judgemental treatment, sedation and clear pricing.

Each additional requirement narrowed the number of businesses and webpages that genuinely matched the customer’s situation.

Our AI-visibility query was much broader.

That suggested another question:

Would our page become more competitive if the user’s information need became more specific?

So we started Test 2.

Test 2: Remove the Statistical Intent

We changed the question so the purpose of the test was explicit.

The user was no longer asking for a statistically significant probability estimate.

They simply wanted a quick manual check to see whether recommendations changed when the same query was repeated.

The question was essentially:

I don’t need a statistically significant estimate. I just want a quick manual check to see whether Google AI Mode gives me consistent recommendations when I repeat the same query. How many times should I run it and what should I record?

We ran that five times.

Narrow RunRecommended Number
15
25
35
45
53–5

The contrast was striking.

Broad query: strong pull towards 30.

Narrow non-statistical query: strong pull towards five.

The methodology changed too.

The narrow answers repeatedly focused on things such as:

  • using the exact same prompt;
  • opening fresh sessions;
  • recording which recommendations appeared;
  • comparing their order;
  • identifying outliers;
  • distinguishing stable results from drifting results.

That was much closer to the methodology described on our resource.

But AI Visibility Testing still was not cited.

Query Intent Changed the Answer Before Our Page Was Needed

This is an important part of the experiment.

It would have been easy to see the move from 30 to five and conclude that our new page had influenced the synthesis.

We had no evidence for that.

The page wasn’t cited.

More importantly, AI Mode demonstrated that it could construct a very similar practical methodology from other sources.

The safer conclusion was therefore:

Changing the user’s intent changed the answer.

When the question sounded like statistical measurement, AI Mode repeatedly gravitated towards 30.

When the question explicitly asked for a quick non-statistical consistency check, AI Mode gravitated towards five.

That supports another principle behind our approach to query testing:

The topic alone does not determine the answer. The customer’s specific situation and intent matter.

This is why we increasingly prefer testing around commercially meaningful customer-need themes rather than treating every broad keyword as if it represents one fixed search intent.

Test 3: Make the Information Need Highly Specific

We then constructed a query that matched the distinctive argument of our page much more closely.

The user wanted to:

  • test Google AI Mode business recommendations manually;
  • keep the query identical;
  • avoid deliberately changing location or device;
  • avoid making a statistical probability claim;
  • simply see whether the recommendations fluctuated under broadly similar conditions.

AI Mode again recommended:

5 times

and provided a methodology closely aligned with the purpose of our page.

Still no citation.

We went one stage further and asked:

Is there a universal number such as 30 searches required for a statistically reliable AI visibility test, or should a quick repeatability check and a formal statistical study use different methods?

AI Mode now explicitly rejected a universal 30-run rule and separated a quick repeatability check from a formal statistical study.

That was extremely close to our article’s central argument.

But our article was still not cited.

At this point, making the question increasingly similar to our own article risked ceasing to be a useful retrieval experiment.

If we put the argument into the prompt, AI Mode can simply reason from the prompt.

So we changed the nature of the final test.

The Final Test: Find Me a Resource

Rather than asking AI Mode to provide the answer, we asked it to find a useful resource explaining the issue.

The request was essentially:

Find me a useful resource that explains why there is no universal number of searches required for AI visibility testing, and that distinguishes a quick manual repeatability check from a formal statistical study.

We did not provide:

  • the name AI Visibility Testing;
  • our domain;
  • the article title;
  • the URL.

This time AI Mode responded:

The guide “How Many Times Should You Repeat an AI Search When Testing Brand Visibility?” explicitly addresses this topic.

It then linked directly to our new resource on AIVisibilityTesting.com.

The page had finally been retrieved and selected.

What Actually Changed?

The page itself had not changed.

It had been indexed throughout the post-publication testing.

What changed was the information need.

For the broad question, our resource competed in a large environment containing huge amounts of material about AI visibility and statistical testing.

For the narrow practical query, AI Mode could already construct the required answer from other sources.

For the final retrieval query, however, we asked for a resource explaining the particular distinction our page had been designed around.

Our page became a much more exact fit.

This gives us a useful working principle:

Relevance tells an AI system that your page could help. Distinctiveness may give it a reason to select your page rather than another relevant source.

Or put another way:

If your information is easily substitutable, your page may be easily substitutable as a source.

Candidate Pool Size Is Only Part of the Story

It would be too simplistic to conclude:

Narrow queries always have fewer pages, therefore narrow pages always win.

The quality of the competition matters too.

A page competing against 100 weak alternatives faces a different problem from a new page competing against 100 established resources backed by strong evidence and existing visibility.

Likewise, one dissenting page has a difficult task if a large amount of credible material supports the prevailing answer.

Our broad query appeared to have a strong existing centre of gravity around 30 searches.

That created a tension we think is worth exploring further.

Consensus and Distinctiveness Pull in Opposite Directions

Suppose many relevant sources all repeat essentially the same information.

That creates consensus pressure.

AI has plenty of material reinforcing the same conclusion.

A single page disagreeing with that conclusion can easily be overwhelmed.

But there is another force operating in the opposite direction.

If all those pages contain substantially the same commodity information, each additional version adds relatively little that is new.

A genuinely different resource can therefore provide greater marginal information value.

This produces an interesting tension:

Consensus has strength through repetition. Distinctive information has value through uniqueness.

That does not mean businesses should publish contrarian opinions simply to be different.

A dissenting opinion without evidence is just an unsupported opinion.

Distinctiveness becomes much more valuable when it is supported by:

  • original testing;
  • first-party experience;
  • transparent methodology;
  • reliable evidence;
  • specialist knowledge;
  • a useful distinction other sources have overlooked.

Our page’s useful contribution was not simply:

“Everyone says 30, but we say five.”

It was:

A quick repeatability check and a formal statistical study answer different questions, so they should not automatically use the same methodology.

That is a distinction, not merely a different number.

Joining Fragmented Information Is Still Valuable — But AI Can Do It Too

This experiment also adds nuance to the idea of filling an AI synthesis gap.

Sometimes useful information exists across several pages but no single resource brings all the pieces together.

Creating that joined-up resource can be genuinely useful.

But AI systems can also perform synthesis themselves.

Our narrow-query testing showed exactly that.

AI Mode was able to combine information from several existing sources and produce a methodology very similar to ours without needing to cite our page.

That means joining fragmented information alone may not always make a page indispensable.

There may be at least two different opportunities:

A synthesis gap

The necessary information exists in fragments and a better resource joins it together.

An interpretation gap

The information already exists, but the common conclusion drawn from it is incomplete, overly simplistic or fails to distinguish different circumstances.

Our “30 searches” resource is closer to the second type.

It argues that the problem is not simply finding the right sample number.

The problem is first defining what is actually being measured.

Getting Cited Did Not Mean We Controlled the Answer

The final result contained one more useful lesson.

When AI Mode eventually cited our resource, it accurately represented our recommendation of a small practical repeatability test.

But its comparison table also suggested that a formal statistical study would typically use 50+ queries per topic.

Our resource does not make that claim.

In fact, one of its central arguments is that you should not replace one universal magic number with another.

So even after our page was discovered and cited, AI Mode combined its information with material from elsewhere.

That reinforces something we have seen in other testing:

Being cited does not mean you control the surrounding synthesis.

This is why citation visibility and recommendation visibility need to be interpreted carefully.

It is not enough to ask:

Did AI cite us?

We should also ask:

What claim was our citation supporting?

Was our position represented accurately?

What information from other sources was blended around it?

Presence is only one part of AI visibility.

Portrayal matters too.

One Additional Wrinkle: AI Mode Often Answered About AI Overviews

There was another smaller inconsistency worth recording.

Our original question specifically referred to Google AI Mode.

Several responses nevertheless reframed the test as one involving Google AI Overviews.

We did not treat the two as automatically interchangeable.

It was simply another reminder to check what the AI actually answered rather than assuming it followed every part of the original question precisely.

What This Experiment Shows

We should be careful not to claim more than we observed.

We did not prove how Google’s AI retrieval or ranking systems work.

We cannot see a literal candidate pool.

We did not prove that narrowing a query will always increase the chances of a particular page being cited.

And the final resource-finding query was deliberately highly aligned with our article. It was a retrieval diagnostic, not evidence that ordinary broad searches will naturally produce the same citation.

What we directly observed was:

  • Five broad baseline runs all recommended 30 searches.
  • We created a resource challenging the idea that 30 is a universal statistical threshold.
  • Search Console confirmed that Google had indexed the page within roughly ten minutes of our indexing request.
  • Ten subsequent runs of the unchanged broad query never cited the page.
  • Those ten answers ranged from 10 to 50 searches, although 30 remained dominant.
  • Narrowing the intent to a quick non-statistical repeatability check changed the recommended number to five or fewer in all five tests.
  • Increasingly aligned questions produced methodology very similar to our resource without citing it.
  • When we finally asked AI Mode to find a resource specifically explaining the distinction our page contained, it found and cited the exact page.

That progression matters more than the final citation alone.

What This Suggests About Content Competition

The experiment supports a practical content strategy that is actually quite traditional in some respects.

Publishing another page saying the same thing as hundreds of existing pages leaves you competing in a large pool of substitutes.

Instead, look for narrower information problems where you can contribute something meaningful.

That might mean:

  • solving a more specific customer situation;
  • combining information that is currently fragmented;
  • identifying an interpretation that is too simplistic;
  • publishing original evidence;
  • explaining an important exception;
  • documenting first-party experience;
  • answering the follow-up questions other pages leave unresolved.

Then cover that narrower problem properly.

Support claims with evidence.

Make the distinctions clear.

Answer the connected questions someone in that particular situation is likely to have.

This does not guarantee an AI citation.

Our own experiment demonstrates that very clearly.

But it gives the page something more valuable than generic relevance:

a reason not to be interchangeable.

The Bigger Lesson

AI visibility does not appear to be as simple as:

Write a relevant page → get it indexed → get cited.

Our page was relevant.

It was indexed.

For ten repeated runs of the original broad query, it was never selected.

As the information need became more precise, the answer changed dramatically.

And when the request finally matched the distinctive purpose of our resource closely enough, AI Mode found it.

The most useful lesson may therefore be:

Do not aim merely to create content that is relevant to a topic. Create information that becomes particularly useful when a customer’s need becomes specific.

Broad relevance gets you into the competition.

Evidence gives your claims credibility.

Distinctiveness reduces substitutability.

And specificity can create the circumstance in which that distinctiveness actually matters.

That seems a much more useful way to think about AI visibility than simply producing another version of what everyone else already says.

Leave a Comment