One of the hardest questions in AI visibility testing comes before you run a single test:
Which prompts should you actually monitor?
It is easy to type a few obvious questions into ChatGPT, Google AI Mode or another AI search platform and start recording which businesses appear.
But a handful of obvious prompts may tell you very little about how your business is represented across the different situations in which a potential customer might actually be choosing a provider.
During the seven-day trial that we covered in our OtterlyAI review, Otterly sent us an onboarding resource called its AI Visibility Prompt Framework.
The framework proposes building a structured set of 100 prompts for a product or service, split between questions about Market Players & Leading Brands and Products Based on Features & Capabilities.
What interested us most was not simply the number 100.
It was the structure Otterly provides to stop those 100 prompts becoming slightly different versions of essentially the same question.
That makes the resource a useful example of something we think matters enormously in AI visibility testing:
Good monitoring starts with good questions.
A Framework Gives Businesses Somewhere Sensible to Start
Someone new to AI visibility monitoring can easily understand the idea of tracking prompts.
Actually deciding what those prompts should be is harder.
Otterly addresses that by spreading its suggested prompt portfolio across different commercial dimensions.
Its first group covers areas including:
- general market leaders;
- affordability and value;
- customer personas and segments;
- regional and local providers;
- reputation and trust.
Its second group examines areas such as:
- quality and performance;
- support and reliability;
- price, value and flexibility;
- ease of use;
- situation-specific requirements.
It also recommends concentrating on mid-to-bottom-funnel searches, where someone already understands the type of product or service they need and is evaluating options or moving towards a decision.
For someone opening an AI monitoring platform for the first time, that is useful guidance.
Instead of tracking 20 variations of:
best accountant
you begin thinking about price, location, customer type, specialist requirements, reputation and particular buying situations.
The result should be a much broader view of how AI systems perceive the market.
Commercial Intent Matters More Than Simply Accumulating Mentions
The emphasis on decision-stage searches particularly fits with something we have repeatedly seen in our own testing.
Consider these two questions:
What are management accounts?
and:
Which Birmingham accountants provide monthly management accounts for manufacturing businesses?
A business could be highly visible for the first question and absent from the second.
Yet the second represents a very different commercial situation.
Someone asking it may already have a business, know the service they require and be trying to decide who can provide it.
That is why we think there is an important distinction between AI visibility generally and AI visibility where it could realistically influence a commercial decision.
The objective is not simply to accumulate mentions.
A more useful question is:
Does the business appear when someone describes a need that it is genuinely well placed to satisfy?
Broad Prompts Still Tell Us Something Important
This does not mean broad prompts are unimportant.
Otterly deliberately includes searches designed to identify general market leaders.
That makes sense.
A question such as:
best accounting software
can help reveal which brands an AI system appears to regard as obvious or established players in the category.
We have seen something very similar in our own testing.
In an experiment across five software markets, broad Google AI Mode searches repeatedly produced small groups of familiar brands. More detailed requirements produced different candidate groups.
We explored that in Does Google AI Mode Default to Market Leaders? What 25 Broad Searches Revealed.
The important point is not that broad prompts are inferior.
They measure something different.
Broad prompts can tell you about general category prominence.
Specific prompts can tell you more about suitability for particular customer needs.
A useful monitoring portfolio may need both.
Situation-Specific Prompts Are Where Things Become Particularly Interesting
One part of Otterly’s framework that particularly caught our attention was its section on situation-specific capabilities.
These prompts are based around particular customer contexts, use cases or requirements rather than general popularity.
That closely resembles the direction our own experiments have been taking.
Suppose an accountancy firm wants manufacturing clients.
You could test:
Which Birmingham accountants specialise in manufacturing?
You might separately test:
Which Birmingham accountants work with Xero?
And perhaps:
Which Birmingham accountants offer fixed monthly fees?
Each question tells you something useful about an individual association.
But a real customer may care about all three at once.
Our experiments have therefore made us increasingly interested in what happens when several commercially important requirements are combined into one realistic customer scenario.
For example:
I run a small manufacturing company in Birmingham with 25 employees. I need an accountancy firm to handle year-end accounts, corporation tax, payroll and monthly management accounts. I use Xero and would prefer a firm offering fixed monthly fees. Which three Birmingham accountancy firms should I consider, and why?
The AI now has to consider several requirements together:
- Birmingham;
- manufacturing;
- company size;
- Xero;
- payroll;
- corporation tax;
- year-end accounts;
- monthly management accounts;
- fixed monthly pricing.
That can create a very different recommendation problem from simply asking for the most prominent accountants in Birmingham.
Detailed Customer Scenarios Also Connect With Query Fan-Out
This becomes particularly interesting when considered alongside query fan-out.
We recently explored this after reading an article by Emily Matthews at NOLA Marketing and comparing NOLA’s explanation with our own experiments.
As we discuss in What Query Fan-Out Means for AI Visibility Testing, one detailed customer question may contain several underlying information needs.
Our Birmingham manufacturing query might require an AI system to find information relating to:
- location;
- manufacturing experience;
- Xero;
- payroll;
- management accounts;
- company size;
- pricing.
The customer sees one question.
The AI system may have to assemble evidence relating to several parts of it before producing an answer.
That gives businesses another reason to think beyond one broad keyword or prompt.
The real commercial unit of interest may sometimes be the customer problem behind the prompt.
Building the Initial Prompt Set Is Only the Beginning
This is where we think AI visibility monitoring becomes an evolving task.
A business might put considerable thought into building its first set of 100 prompts.
Some will probably prove extremely useful.
Others may initially look commercially promising but become less interesting once the monitoring begins.
Perhaps a prompt repeatedly produces very generic results.
Perhaps two prompts turn out to measure essentially the same thing.
Perhaps a customer segment proves less commercially important than expected.
Or perhaps one apparently modest prompt reveals an interesting recommendation pattern and deserves much deeper investigation.
The original prompt research has not failed.
The testing has taught you something.
And that information should help improve the testing itself.
Your First Prompt Portfolio Is a Prospecting Map
Gold prospecting provides a useful analogy.
A prospector does not necessarily know where the valuable deposits are before exploration begins.
They identify promising ground based on the information available and investigate it.
Some areas justify deeper exploration.
Others eventually prove less worthwhile and attention moves elsewhere.
Your first AI prompt portfolio is similar.
It is a prospecting map, not a list of proven gold deposits.
A query might appear commercially valuable when you first add it to the monitoring set.
Only after observing the answers over time might you discover whether it is genuinely giving you useful information.
And occasionally you find promising ground.
Suppose this query produces a particularly interesting pattern:
best accountant for manufacturing businesses
That may encourage you to explore nearby questions:
accountant for manufacturers using Xero
accountant for growing manufacturing companies
accountant for manufacturers needing monthly management accounts
accountant for manufacturers wanting fixed monthly fees
One prompt has now helped expose a broader commercial theme worth investigating.
The important point is that prompt refinement should not be seen as correcting a failed initial list.
Learning which questions are worth monitoring is itself one of the outputs of monitoring.
A Prompt Portfolio Needs Stability and Exploration
There is an obvious tension here.
If you continually replace every prompt, it becomes difficult to measure changes over time.
But if you never change anything, your monitoring set can gradually become disconnected from what actually matters to customers and the business.
One practical way to think about this is to maintain two groups of prompts.
Core Prompts
These represent important customer needs that remain commercially relevant.
Keeping them relatively stable gives you continuity and allows visibility to be compared over time.
Exploratory Prompts
These allow you to investigate:
- new customer situations;
- emerging needs;
- wording discovered through customer conversations;
- questions appearing in Search Console;
- new products or services;
- interesting patterns discovered in previous tests.
Some exploratory prompts will prove unimportant and disappear.
Others may become useful enough to join the core set.
The process therefore becomes:
research → choose → monitor → learn → refine → monitor again
That feels much closer to commercial reality than creating a list once and assuming prompt research is finished.
Otterly Encourages Businesses to Look Beyond AI-Generated Prompt Lists
Another part of the framework we particularly liked is that Otterly does not suggest simply asking an AI model to invent 100 prompts and stopping there.
Its resource recommends several potential inputs, including its own prompt research tools, Google Search Console, Copilot, Google searches, ChatGPT and customer surveys.
The customer research element is especially useful.
Otterly suggests asking existing customers what they searched for, what mattered when choosing a provider and which features or capabilities they considered.
That keeps prompt research tied to actual commercial behaviour.
Google Search Console can help for a similar reason.
A Search Console query does not prove somebody entered exactly the same wording into ChatGPT or Google AI Mode.
But it does show that a real searcher expressed that information need.
That can provide valuable raw material for deciding what customer themes are worth testing in AI search.
Why 100 Prompts Can Be a Useful Starting Framework
We do not think the most important question is whether 100 happens to be the perfect number of prompts for every business.
The practical value of Otterly’s framework is the breadth and structure it encourages.
The document itself explains that the standard framework is intended to create consistency and comparability across markets and business units.
In the context of helping someone get started with an emerging form of marketing measurement, that makes sense.
A structured 100-prompt portfolio can prevent a business from monitoring an arbitrary handful of closely related searches and assuming the resulting visibility percentage represents the whole market.
We take a similar practical approach with manual testing.
When we run an exploratory query five times, we are not claiming that five searches establish statistical certainty.
We are trying to learn more than one search can tell us while keeping the test manageable for an ordinary business or website owner.
There is an important difference between:
a practical framework that helps a business make progress
and:
a universal statistical rule that everyone must follow.
For most businesses, the commercial question is whether the testing produces sufficiently useful information to make better decisions.
Choosing the Right Prompts Solves Only Half the Problem
A carefully constructed prompt portfolio answers one major question:
What should we test?
There is another:
How often should we test it?
AI-generated answers can change when exactly the same question is repeated.
Businesses may appear or disappear.
Recommendation order can change.
Different supporting websites may be cited.
That is why we separately explored How Many Times Should You Repeat an AI Search When Testing Brand Visibility?.
A broad prompt portfolio gives you coverage across customer needs.
Repeated observations give you a better understanding of variation within those needs.
Both matter.
Choose Intelligently, Observe What Happens, Refine Intelligently
The biggest lesson we took from OtterlyAI’s 100-prompt framework was not that every business needs exactly the same permanent list of 100 questions.
It was that prompt selection deserves much more thought than choosing a few obvious search phrases and looking to see whether your business appears.
A useful process might instead have three stages.
1. Choose Intelligently
Start with a structured collection of commercially meaningful questions.
Include broad market searches, customer personas, features, locations and specific situations rather than concentrating everything around one generic term.
2. Observe What Happens
Monitor the prompts.
Repeat important searches.
Look at which businesses appear, how consistently they appear and what types of customer requirements seem to change the recommendations.
3. Refine Intelligently
Keep commercially important core prompts so you can observe trends.
But continue exploring.
Replace questions that prove weak. Develop promising themes. Add new customer needs as you discover them.
Your monitoring should become better informed by the information the monitoring itself produces.
Our Experience With OtterlyAI Has Reinforced Our Positive View of the Platform
We have already explained in our full OtterlyAI review why we are comfortable recommending the platform for businesses, website owners and agencies that reach the point where manual AI visibility monitoring becomes impractical.
This onboarding resource gives us another reason to feel positive about that recommendation.
Otterly is not simply giving users software that runs searches and presents visibility data.
It is also trying to help people understand what they should be monitoring in the first place.
That matters.
AI visibility monitoring is still an emerging area. Many businesses are likely to understand that they should be checking platforms such as ChatGPT and Google AI search without yet knowing which questions are worth tracking, how broad their test should be or what they should do with the results.
A practical framework helps turn that uncertainty into something businesses can actually work with.
For a small number of commercially important prompts, manual testing remains extremely useful because it lets you inspect the answers closely and understand what is happening.
But as the prompt portfolio expands across different customer needs, platforms and dates, consistent manual monitoring quickly becomes a substantial task.
That is where a platform such as OtterlyAI starts to make much more commercial sense.
The objective is not to create a supposedly perfect prompt list on day one.
It is to start with a sensible framework, observe what the results reveal and gradually develop a monitoring portfolio that better reflects the customers, situations and commercial decisions that actually matter to the business.
Prompt research is not something you complete before AI visibility monitoring starts. It is something the monitoring itself should continuously help you improve.