How to Build an AI Visibility Query Set Around Real Customer Situations

Choosing the right searches is one of the most important parts of AI visibility testing.

But there is a difference between creating a long list of prompts and creating a query set that genuinely represents potential customers.

You could test:

  • best accountant in Bristol
  • good accountant in Bristol
  • recommended accountant in Bristol
  • top accountant in Bristol

That gives you four prompts.

But they are essentially variations of the same broad customer situation.

Our recent experiments suggest a more useful approach:

Build your AI visibility query set around the circumstances that could genuinely change which business or product is the right fit.

That distinction matters because prompt variation is not the same as customer-situation coverage.

In an earlier experiment, I looked at how choosing more commercially meaningful AI search prompts changed the products recommended.

Since then, further tests with accountants, commercial cleaners and recruitment CRM software have helped turn that idea into a more structured method.

Why Broad Queries Are Still Useful

Broad searches should not disappear from an AI visibility test.

They provide a benchmark.

For example, we asked Google AI Mode three times:

Can you recommend CRM software for a small business?

The results were extremely stable.

HubSpot CRM, Pipedrive, Zoho CRM and Capsule CRM appeared in all three runs.

Given how little information the query supplied about the customer, a group of familiar general-purpose CRM platforms was a reasonable result.

But then we changed the customer.

Can you recommend CRM software for a small recruitment agency?

The recommendation set changed substantially.

Recruit CRM, Giig Hire, Loxo and Zoho Recruit appeared consistently.

HubSpot and Pipedrive disappeared.

The underlying product category had not changed.

The customer had.

And that changed what a suitable product needed to do.

That is the type of difference a useful AI visibility query set should capture.

We Saw the Same Pattern With Accountants

We also asked Google AI Mode:

Can you recommend an accountant for a small limited company in Bristol?

Across three identical runs, nine different accountancy firms appeared.

We then specified:

a small construction limited company in Bristol

The recommendation set became much more consistent.

Then we added another detail:

the company regularly uses subcontractors

The recommendations changed again.

EasyAccounts, which had not appeared in either of the previous query groups, appeared in all three subcontractor searches.

Only afterwards did we inspect its website.

We found detailed information about:

  • construction businesses;
  • CIS contractors and subcontractors;
  • subcontractor verification;
  • monthly CIS returns;
  • CIS deductions;
  • payroll.

That sequence was important.

We chose the customer situation first, ran the searches and then investigated the businesses that appeared.

We did not find an interesting website and construct a prompt specifically designed to match it.

But Not Every Additional Detail Produced the Same Effect

This led to another question.

Was AI changing its recommendations simply because we were making the prompt more detailed?

We tested that with recruitment CRM software.

First:

Can you recommend CRM software for a small recruitment agency?

Then:

Can you recommend CRM software for a small recruitment agency that mainly places temporary staff?

Adding temporary staffing materially changed what the software needed to do.

AI Mode started prioritising things such as:

  • worker availability;
  • shift scheduling;
  • timesheets;
  • compliance;
  • payroll;
  • pay-and-bill.

The recommendation set changed accordingly.

Vincere, for example, appeared in all three temporary-staffing runs after appearing in none of the general recruitment-agency runs.

We then tried a different extra detail:

Can you recommend CRM software for a small recruitment agency that has been trading for five years?

This also made the query more specific.

AI adapted its explanation, talking more about established databases, automation, migration and business development.

But the main recommendation set remained broadly similar to the original recruitment-agency search.

That gave us a more useful working hypothesis:

The kind of specificity may matter more than the amount of specificity.

Details that materially change what a customer needs may be much more valuable to test than details that merely make the description longer.

Build the Query Set Around Six Questions

From those experiments, I now think a practical query set can be built around six simple questions.

1. What Does the Customer Need?

Start with the basic product or service.

For example:

Can you recommend an accountant for a small limited company in Bristol?

or:

Can you recommend CRM software for a small business?

This becomes your broad benchmark.

It shows what AI recommends when relatively little is known about the customer.

2. Who Is the Customer?

Now identify the type of customer.

For example:

Accountant for a construction company

CRM for a recruitment agency

Commercial cleaner for a dental practice

Different customer groups can introduce very different requirements.

A dental practice and an ordinary office may both need a cleaner, but the factors influencing provider suitability are not necessarily the same.

Likewise, a recruitment agency and a general small business may both need CRM software, while requiring very different functionality.

3. How Do They Operate?

This was one of the strongest dimensions in our experiments.

Ask what the customer actually does that could affect which provider is suitable.

Examples include:

Construction company regularly using subcontractors

Recruitment agency mainly placing temporary staff

Business operating from several locations

Company selling internationally

These circumstances may introduce requirements that were absent from the broader search.

4. What Specific Problem Are They Trying to Solve?

Customers often search because something has happened.

That can create much stronger buying intent than simply searching for a category of provider.

For an accountant, examples might include:

Company with an overdrawn director’s loan account

Business struggling with CIS returns

Company needing to change accountants

For software:

Recruitment agency needing digital timesheets

Business needing CRM software that integrates with its existing systems

For a service business:

Dental practice needing cleaning outside normal opening hours

These are not just descriptive details.

They help explain why the customer is looking now.

5. What Could Rule a Provider In or Out?

The next category is the customer’s practical constraints.

These might include:

  • location;
  • budget;
  • fixed pricing;
  • opening hours;
  • turnaround;
  • software compatibility;
  • integrations;
  • implementation time;
  • minimum contract;
  • availability.

These details can be highly commercially relevant because they may eliminate otherwise suitable providers.

A cleaner might offer exactly the right service but be unavailable when the premises are closed.

A software product might have ideal functionality but be outside the buyer’s budget.

An accountant might specialise in the right industry but not support the customer’s existing accounting platform.

Those are real selection criteria.

6. Does the Extra Detail Actually Matter?

Occasionally add a lower-relevance comparison.

Our recruitment CRM test used:

has been trading for five years

The detail was real and allowed AI to infer some additional needs.

But it did not redefine the required CRM functionality in the way that temporary staffing did.

A control like this can help distinguish:

AI changed the recommendations because the customer genuinely required something different

from:

AI changed the recommendations simply because I made the prompt longer.

You do not need a control for every query.

But occasionally using one can make an experiment much more informative.

Turn Those Six Questions Into a Customer-Situation Matrix

Suppose we were testing an accountancy firm in Bristol.

Instead of producing dozens of variations of “accountant Bristol”, we could create something like this:

DimensionExample query
Broad needAccountant for a small limited company in Bristol
Customer typeAccountant for a small construction company in Bristol
Operating modelAccountant for a construction company regularly using subcontractors
Specific problemAccountant for a company with an overdrawn director’s loan account
Software requirementAccountant for a construction company using Xero
Service requirementAccountant handling both payroll and CIS
Buying constraintAccountant offering fixed monthly fees
ControlAccountant for a company that has traded for five years

These are not necessarily the exact queries an accountancy firm should use.

The point is the structure.

Each search should test a genuinely different reason why one accountant might be more appropriate than another.

Prompt Variation Can Still Be Tested — Just Don’t Confuse It With Customer Coverage

There is nothing wrong with testing different wording.

For example:

best accountant in Bristol

and:

can you recommend an accountant in Bristol?

might produce different AI responses.

That is worth investigating if your objective is to understand how sensitive an AI system is to wording.

But that is a different experiment.

It is prompt variation.

By contrast:

accountant for a construction company using subcontractors

tests a different customer situation.

A useful AI visibility programme may eventually test both.

But putting 100 slightly different phrases into a monitoring platform should not create the impression that you have covered 100 different customer needs.

How Many Customer Situations Should You Test?

There is no universal number.

For a small manual experiment, 10–20 genuinely different customer situations can be a useful starting point.

But the number matters less than whether each query tests something distinct.

Ask of every prompt:

What does this query tell me that the others do not?

If the answer is simply that it uses a different adjective, it may not deserve a separate place in the core query set.

A narrow specialist business may need fewer queries.

A company with several products, customer types or operating models may need many more.

The aim is coverage, not volume for its own sake.

Repeat the Important Queries

Choosing meaningful queries is only one part of the test.

AI-generated recommendations can vary between identical searches.

In my earlier repeated testing, the same Google AI Mode product query produced different shortlists even though the prompt itself did not change.

That is why a repeatable manual AI visibility test needs several runs rather than one favourable or unfavourable result.

For an initial manual test, three repeated runs can expose obvious variability.

Queries that matter particularly strongly to the business may justify more.

The important thing is to keep the query unchanged between repetitions when you are measuring variability.

Otherwise you are changing two things at once.

Don’t Record Everything as the Same Type of Visibility

Once the queries are chosen, another distinction becomes important.

AI can make a business visible in several different ways.

It can:

  • recommend the business;
  • mention the brand;
  • cite the company’s website;
  • describe the business accurately or inaccurately.

Those are not necessarily the same outcome.

In one of our earlier experiments, a single AI Mode response contained three recommended products but many more linked websites.

That is why I now separate being mentioned, being cited and being recommended rather than treating every appearance as one generic visibility metric.

For each query, it can therefore be useful to record:

Recommendation
Was the company actually presented as a suitable choice?

Mention
Was the brand named somewhere in the response?

Citation
Was the company’s own website used as a source?

Representation
What did AI actually tell the customer about the company?

The last point matters because a mention only has value if the description gives the customer a reasonably accurate impression of the business.

A simple AI brand-understanding sense check can help explore that separately.

Not Every Query Deserves Equal Importance

Suppose an accountancy firm specialises in construction businesses using subcontractors.

Its results might eventually look something like:

Generic small-business accountant query:
80% recommendation visibility

Construction business using subcontractors:
20% recommendation visibility

A headline average could make the firm’s AI performance look reasonably good.

Commercially, however, the second figure might matter far more.

This is why measuring AI share of voice across different customer requirements can reveal more than one overall percentage.

A simple way to handle this is to label queries:

High importance
Represents a valuable customer and meaningful buying situation.

Medium importance
Relevant, but less central to the business.

Benchmark
Useful for understanding broad-market visibility.

You do not necessarily need a complicated weighted formula.

The important thing is to avoid allowing a large collection of generic searches to hide weak visibility in the customer situations the business most wants to win.

Choose the Query Before Looking for the Winner

There is one methodological rule I would keep wherever possible:

Choose the customer situation before investigating which businesses are likely to match it.

Our EasyAccounts result is a good example.

We did not first discover a page about CIS subcontractors and then construct a search to make EasyAccounts appear.

The sequence was:

Customer situation

→ construction business using subcontractors

AI test

→ EasyAccounts appeared 3/3

Website investigation

→ detailed CIS contractor and subcontractor proposition discovered

That does not prove why AI Mode selected EasyAccounts.

But it makes the observation much more useful than designing the query after seeing the website.

A Practical Query-Building Process

The method can therefore be kept fairly simple.

1. Define the broad product or service.

What is the basic thing the customer wants?

2. Identify the major customer types.

Who genuinely buys it?

3. Map the operational differences.

What circumstances change what those customers need?

4. Identify the problems that trigger a search.

Why might somebody be looking for a provider now?

5. Add meaningful buying constraints.

What could make one apparently suitable provider unsuitable?

6. Keep broad benchmark queries.

They show the general recommendation landscape.

7. Occasionally use a lower-relevance control.

This can help reveal whether the type of detail matters.

8. Remove unnecessary duplicates.

Make sure each core query represents something genuinely different.

9. Repeat the important searches.

A single response is not a reliable visibility measurement.

10. Record recommendations, mentions, citations and representation separately.

They describe different forms of AI visibility.

11. Prioritise commercially important situations.

Visibility matters most where the customer matters most.

The Query Set Defines the Measurement

This may be the most important principle.

There is no meaningful AI visibility percentage without the query set behind it.

If you test mostly broad category searches, you are measuring broad-category visibility.

If you test specialist customer situations, you are measuring specialist visibility.

If you test mainly informational questions, you are measuring something different from queries where someone is actively looking for a provider.

So whenever you see a figure such as:

Our AI share of voice is 35%.

there should be an immediate follow-up question:

35% across which customer situations?

The answer determines what the number actually means.

Final Thought

It is easy to make AI visibility monitoring look sophisticated by tracking hundreds of prompts.

But the number of prompts is not the most important thing.

The real challenge is deciding whether those prompts represent the situations faced by genuine customers.

Who are they?

How do they operate?

What problem are they trying to solve?

What constraints do they face?

And which of those facts would genuinely change which provider is right for them?

That leads to a much more useful question than:

How often does AI mention us?

Instead ask:

When someone like our real customer describes the situation they are actually in, does AI recognise that our business could be the right fit?

That is the visibility worth measuring.

Leave a Comment