I have spent a lot of time manually testing how businesses, products and websites appear in AI-generated search results.
That has involved repeating the same searches, changing customer requirements, comparing competitors, checking citations and running the same question through different AI platforms.
Manual testing is extremely useful because it helps you understand what you are actually measuring.
It also becomes time-consuming very quickly.
That is why I recently signed up for the free trial of OtterlyAI.
Having now used it on AIVisibilityTesting.com, I think it is a tool I can recommend for businesses, website owners and agencies that want to monitor AI visibility more systematically.
Not because it removes the need to think about what you are measuring.
It doesn’t.
What it does is automate much of the repetitive work once you have decided what is worth monitoring.
What Does OtterlyAI Do?
OtterlyAI is an AI search monitoring platform.
You provide—or research—a collection of prompts that potential customers might ask AI systems. Otterly can then monitor those prompts across multiple AI search platforms and record information such as:
- whether your brand appears;
- where it appears;
- which competitors appear;
- whether your website is cited;
- which other domains are being cited;
- how your visibility compares with competitors;
- and how those results change over time.
It also includes prompt research, competitive benchmarking, sentiment analysis and downloadable reporting.
You can see the current feature set on the OtterlyAI website.
For me, however, the value of the tool makes most sense when its features are connected to the practical problems involved in AI visibility testing.
Otterly Suggested My First 15 Prompts
When I entered AIVisibilityTesting.com, Otterly suggested 15 prompts to monitor and also identified competing websites.
That immediately solves a genuine beginner problem.
If you have never done AI visibility testing before, being asked to think of 15 different questions that prospective customers might ask can be surprisingly difficult.
An automatically generated starting set gets you moving.
I would not treat those suggestions as the final prompt set.
In my own case, several of the suggested prompts focused quite heavily on AI visibility tools and platforms. That makes sense from Otterly’s interpretation of the site, but AIVisibilityTesting.com is primarily an independent publication about testing and understanding AI visibility rather than an AI monitoring software product.
Some of the prompts were therefore more useful than others.
But I don’t see that as a reason not to use the feature.
I see it as a starting point.
You can begin with the suggestions, look at the results and gradually replace, refine or expand the prompts as you learn more about what people might realistically ask.
That fits closely with something I have already explored in How to Choose AI Search Prompts That Actually Matter to Your Business.
Automation does not make prompt selection unimportant.
If anything, it makes selecting the right prompts even more important because you may eventually monitor dozens of them.
Your Prompt Set Should Develop Over Time
I don’t think AI visibility monitoring should ever be treated as a one-and-done exercise.
A business may start with 15 prompts.
Some will turn out to be useful.
Some may be too broad.
Others may describe customers who are unlikely to buy anything.
And the first results may reveal entirely new customer situations worth exploring.
That is why I prefer thinking in terms of customer needs rather than trying to discover every possible sentence someone might type into an AI system.
I explored that approach in How to Build an AI Visibility Query Set Around Real Customer Situations.
The Otterly suggestions can therefore be the beginning of the research process rather than its conclusion.
Start with them.
Look at what the AI systems actually return.
Think about why particular competitors appear.
Consider which searches matter commercially.
Add better prompts as you understand your market.
That seems much more realistic than expecting any software to discover the perfect permanent list automatically.
Automated Monitoring Solves the Repetition Problem
This is probably the biggest reason I can see for using a tool like OtterlyAI.
One AI search tells you very little.
I demonstrated that in How to Run a Repeatable Manual AI Visibility Test.
The same prompt can produce different recommendations when repeated.
That means checking once and finding your business—or failing to find it—does not tell you how consistently visible it is.
I have also explored the question in How Many Times Should You Repeat an AI Search When Testing Brand Visibility?.
Manual repetition is perfectly practical when you are investigating a small number of prompts.
The problem comes when you want to monitor:
- 15 prompts;
- several AI platforms;
- multiple competitors;
- different dates;
- and changes over time.
The number of observations soon multiplies.
At that point, automation starts making considerably more sense.
Otterly can repeatedly run the monitoring while retaining the results so that you can look for patterns rather than continually performing the searches yourself.
That is the natural point where I think manual testing and automated monitoring complement each other.
Manual testing helps you understand the measurement. Automated monitoring helps you scale it.
Monitoring More Than One AI Platform Matters
Another useful feature is the ability to monitor multiple AI search platforms.
Our own experiments have already shown why this matters.
In Google AI Mode vs ChatGPT: I Asked the Same Buying Question Five Times, I gave Google AI Mode and ChatGPT exactly the same detailed customer requirement.
The recommendation patterns were not the same.
One accountancy firm ranked first in all five Google AI Mode searches but never ranked first in ChatGPT.
Another ranked first four times in ChatGPT but never first in Google.
So saying:
“We’re visible in AI search.”
isn’t necessarily enough.
A better question may be:
“Where are we visible?”
Monitoring several AI platforms from one place makes that comparison much easier than manually maintaining separate tests.
I Particularly Like the Separation Between Mentions and Citations
One of the first things I noticed in my Otterly report was that it did not simply give me one generic visibility number.
It separately recorded whether my brand was mentioned and whether my domain was cited.
That distinction matters.
In my initial 15-prompt report, AIVisibilityTesting.com received brand appearances for two prompts.
But its own domain was not recorded as cited for any of them.
Those are two different results.
Our experiments have repeatedly found that being mentioned, being recommended and being used as a source should not automatically be treated as the same thing.
I explain the distinction in Mentioned, Cited or Recommended: What Kind of AI Visibility Do You Have?.
I also tested citations directly in I Compared AI Citations With Recommendations Across 10 Google AI Mode Searches.
In that experiment, a website being cited did not necessarily mean the business was one of the companies being recommended.
So I would much rather have software expose these different signals than collapse everything into one apparently impressive visibility score.
Competitor Monitoring Could Be Extremely Useful
Otterly also identified competing brands and showed how often they appeared across the same prompts.
This provides an important additional layer.
Suppose your company appears in 20% of monitored AI answers.
Is that good?
It depends.
If every major competitor appears 80% of the time, your 20% tells one story.
If most competitors rarely appear either, it tells another.
Competitor comparison helps provide context.
It can also reveal something more useful than the headline figure:
Which particular customer situations are competitors winning?
That could lead you back to their websites.
What evidence do they provide?
Do they clearly describe a specialist service?
Do they answer questions you don’t?
Are third-party websites repeatedly being used as sources about them?
The monitoring result becomes the beginning of an investigation.
But Don’t Confuse Visibility With Commercial Value
This is where I think human interpretation remains essential.
Suppose Otterly eventually tells me:
AIVisibilityTesting.com has 70% visibility.
My next question would be:
70% visibility across which prompts?
A business could dominate a collection of broad informational searches that rarely lead to a customer.
Another business might appear less frequently overall but perform extremely well when someone describes precisely the problem that business specialises in solving.
That is why I developed The RAMP Framework.
RAMP asks four questions:
R — Relevance: Was this a commercially useful search to appear for?
A — Appearance: Did the business actually appear?
M — Mention Quality: How was the business presented?
P — Pathway: Was there a practical route from the answer towards the business?
Otterly can help enormously with collecting the Appearance evidence.
It can also provide useful information around mentions, citations, competitors and sentiment.
But businesses should still ask whether the prompts being measured matter in the first place.
A dashboard cannot make an irrelevant query commercially valuable.
You Should Still Read the Actual AI Answers
There is another reason I would not rely exclusively on headline metrics.
AI can mention your business and still misunderstand what you do.
I explored this with the Interpretation Test.
That issue is particularly relevant to AIVisibilityTesting.com because the name itself is also a descriptive phrase.
If an automated report records the words “AI Visibility Testing”, I still want to know whether the underlying AI answer genuinely recognised my website or merely used those words generically.
Metrics can tell you where to look.
You should still occasionally read the underlying answers.
AI visibility is not just about being counted.
It is about understanding why you were counted and what the AI actually said about you.
So, Would I Recommend OtterlyAI?
Based on my initial use of the free trial, yes.
I think OtterlyAI solves a genuine problem.
You can perform AI visibility testing manually—and I actually recommend doing some manual testing first because it teaches you how variable and nuanced the results can be.
But once you want to monitor a meaningful collection of prompts across multiple AI systems and see how those results change over time, manual testing becomes increasingly cumbersome.
That is where Otterly starts to make sense.
I particularly like:
- the automated prompt suggestions as a starting point;
- the ability to add and develop your own prompt set;
- monitoring across multiple AI platforms;
- competitor comparisons;
- separate tracking of brand mentions and website citations;
- historical monitoring rather than one-off checking;
- and the ability to export the data for further analysis.
My main qualification is equally important:
Don’t outsource the thinking to the tool.
Otterly can tell you what happened across the prompts you asked it to monitor.
It cannot make a poor prompt commercially meaningful.
It cannot decide for you which customer situations matter most.
And a visibility percentage still needs interpretation.
Used in that way, I think Otterly fits very naturally into the testing process I have been developing on this site:
Explore manually → identify meaningful customer situations → build a prompt set → automate the repetitive monitoring → investigate what changes.
That isn’t a one-and-done AI visibility audit.
It is an ongoing measurement process.
And that is exactly why a monitoring tool becomes useful.