I recently read an interesting case study published by OtterlyAI about NOLA Marketing’s work for Single Point of Contact.
It caught my attention because several parts of the process closely match conclusions we have reached independently through our own AI visibility testing.
The headline results are certainly impressive.
OtterlyAI reports that, over a three-month period, SPOC’s AI mentions increased from 30 to 160, while citations increased from 665 to 3,617. The company also moved from third to first in AI citation rank.
OtterlyAI also reports that SPOC’s head of sales saw a 30% increase in inbound leads alongside the programme.
But those numbers aren’t actually the part of the case study I found most interesting.
What caught my attention was how NOLA decided what to test, what it did with the results, and how the testing influenced what happened next.
For anyone trying to understand what useful AI visibility testing might look like in practice, I think there are some important lessons here.
A note on the evidence: this is a case study published by OtterlyAI about an agency using OtterlyAI. The SPOC results discussed below are reported by OtterlyAI and NOLA Marketing. I haven’t independently reproduced them, so I am treating this as a useful real-world case study rather than proof that the same approach will produce the same results for every business.
They Started With the Buyer Journey, Not Just a List of Keywords
NOLA eventually built a final set of 64 prompts.
It initially focused on Google AI Overviews, Google AI Mode and ChatGPT, while Copilot became increasingly important because many of SPOC’s customers operate in Microsoft environments.
But the important part was not simply the number of prompts.
NOLA did not just take a list of existing SEO keywords and convert each one into a question.
Instead, it worked through the customer’s buying journey with the client.
Particular attention was given to questions nearer the bottom of that journey — situations where someone might be getting closer to choosing a provider.
That makes a lot of sense.
Consider two searches:
What is a white-label help desk?
and:
What should an MSP look for when choosing a white-label help desk provider?
Both could be relevant to the same company.
But they don’t necessarily represent the same stage of a customer’s journey.
The first person may simply be learning what the service is.
The second may be much closer to considering which company could provide it.
That distinction matters when measuring AI visibility.
It also closely matches something we have discussed in our guide to choosing AI search prompts that actually matter to your business.
A prompt can be highly relevant to the subject of a business without necessarily being commercially important.
A Visibility Percentage Is Only as Useful as the Prompts Behind It
Imagine testing 50 broad informational prompts and finding that your business appears in 35 of them.
You might reasonably report a 70% presence rate.
But what does that actually mean commercially?
If relatively few of those searches represent situations where someone might genuinely consider buying from you, an impressive-looking visibility percentage could have limited practical value.
Conversely, appearing consistently across a smaller group of highly relevant searches could potentially matter much more.
This is consistent with the principle behind our RAMP Framework: visibility matters much more when the underlying search has genuine commercial relevance.
That is why prompt selection matters so much.
The measurement can only tell you what is happening within the questions you decided to measure.
Choose poor prompts and you can produce a very accurate measurement of something that isn’t particularly useful.
Testing Became an Input Into the Content Strategy
Another part of the OtterlyAI case study is particularly interesting.
NOLA wasn’t simply creating content and then using monitoring software afterwards to report whether visibility had increased.
The monitoring influenced what happened next.
Prompt research established the kinds of buyer questions worth tracking.
Ongoing monitoring then revealed gaps where competitors were appearing and SPOC was not.
Those findings helped shape future content.
The process effectively became:
Choose relevant prompts → measure visibility → identify gaps → create useful content → measure again → adjust.
That is a much more interesting use of AI visibility testing.
The testing isn’t the objective.
It provides information that can help someone make better decisions.
The New Content Was Built Around Buyer Questions
The content subsequently created for SPOC is also worth examining.
OtterlyAI says NOLA built new resource-centre hubs around SPOC’s core service categories, containing content designed to answer questions buyers were already putting to AI engines about white-label IT, help desk and SOC services.
According to the case study, by June 2,657 of SPOC’s 3,617 citations — approximately 73% — came from the new question-and-answer content created for the programme.
That is an interesting result.
But I don’t think the correct conclusion is:
AI likes FAQ pages, so everyone should create lots of them.
The case study doesn’t prove that.
What interests me much more is the relationship between the questions being researched and monitored and the content subsequently created to address them.
Prompt research identified the types of questions buyers were asking.
Monitoring revealed where SPOC was absent but competitors were appearing.
Those findings helped influence the content created next.
The useful lesson is not necessarily the format of the page.
It is that the content was being developed around identifiable information needs rather than simply around a list of services the company wanted to describe.
Mentions and Citations Can Tell Different Stories
The OtterlyAI case study also provides a good example of why AI visibility cannot always be reduced to one number.
A later three-month snapshot showed that, against SPOC’s tracked competitive set, it accounted for roughly 73% of brand mentions.
But when Otterly looked at all domains being cited by AI engines within the tracked category, SPOC’s own domain accounted for about 9% of citations.
Both figures can be correct.
They simply describe different aspects of visibility.
One asks something similar to:
How often are we appearing compared with these competitors?
The other asks:
What proportion of the wider cited source landscape comes from our own website?
That distinction matters.
It also fits with something we have found in our own testing: being mentioned, cited and recommended are different forms of AI visibility.
A citation, for example, does not necessarily mean that the business itself is being recommended.
That is why an impressive-looking AI visibility metric needs context.
A #1 citation rank, a high Share of Voice or a large percentage increase can all be useful measurements.
None tells the whole story by itself.
This Is Where Automated Monitoring Starts to Make Sense
There is another practical lesson hidden inside the scale of the NOLA programme.
NOLA was monitoring 64 prompts across multiple AI engines.
Even a single full manual pass at that scale can mean reviewing hundreds of generated responses.
Repeat that regularly over weeks or months and the workload quickly becomes difficult to sustain.
That doesn’t mean every business needs automated AI visibility software from day one.
For somebody initially trying to understand how their business appears in AI search, a much smaller manual test can still be useful.
But dozens of prompts across several platforms, monitored continually over time, is a different proposition.
That is where automation starts solving a very obvious practical problem.
It is also one of the reasons I concluded in my separate OtterlyAI review that it is a tool I can recommend for ongoing monitoring once manual checking becomes impractical.
The important point is that automation does not make poor prompts valuable.
It simply makes useful testing easier to repeat at scale.
What About the 30% Increase in Leads?
Potentially the most commercially important figure in the whole case study is not the citation growth.
It is the reported 30% increase in inbound leads alongside the programme.
But this needs to be treated carefully.
OtterlyAI itself does not claim that improved AI visibility definitely caused the entire increase.
The two happened during the same period.
Other marketing activity, seasonal changes, sales activity or external factors could also have influenced enquiries.
That qualification matters.
If citations increase from 665 to 3,617, that tells us citations increased within the measurement being used.
It doesn’t automatically prove those additional citations produced customers.
But businesses ultimately are not trying to win an AI visibility competition.
They want outcomes.
More enquiries.
More customers.
More sales.
More registrations.
Or whatever matters to that particular organisation.
That is why the reported increase in inbound leads is interesting even though it cannot be attributed solely to AI visibility.
Mentions and citations are useful measurements.
The eventual question for a business is whether increased visibility contributes to commercially useful outcomes.
What I Think This Case Study Really Demonstrates
For me, the biggest lesson from the OtterlyAI/NOLA Marketing case study isn’t:
Create question pages and your AI citations will increase.
The case study doesn’t prove that.
It also isn’t:
Achieve a higher citation rank and your leads will increase by 30%.
It doesn’t prove that either.
The more useful lesson is the process.
NOLA:
- worked through the buyer journey;
- selected prompts around real buyer questions;
- included questions closer to a commercial decision;
- measured where SPOC and its competitors were appearing;
- identified gaps;
- used those findings to influence new content;
- continued monitoring after publication;
- used the resulting data to inform what happened next.
That is a feedback loop.
And I think that is one of the most useful ways of thinking about AI visibility testing.
Don’t Test Just to Produce a Score
There is a temptation with any new marketing measurement to reduce everything to one number.
Your AI Visibility Score is 42.
Your Share of Voice is 18%.
You rank third.
Those numbers can be useful.
But the more interesting questions are underneath them.
Which prompts are you appearing for?
Are those prompts commercially important?
Which important prompts are you missing from?
Who appears instead?
Are you being mentioned, cited or actually recommended?
Which pages and sources are AI systems relying on?
Does your website adequately answer the questions potential customers are asking?
Does anything change after you improve that information?
And ultimately, does any of this appear to contribute to useful business outcomes?
Those are much harder questions to compress into one attractive dashboard number.
But they are probably the questions that matter.
The original OtterlyAI case study doesn’t give us a universal formula for improving AI visibility.
No single case study could.
What it does provide is a useful real-world example of AI visibility measurement being used as part of an ongoing decision-making process.
For me, that is considerably more interesting than simply moving from third place to first.