It's easy to check whether AI recommends your business. Open ChatGPT, type a question, read the answer. It takes thirty seconds and it feels definitive. It's also, on its own, one of the least reliable ways to test this, and it's worth explaining exactly why, because the gap between a quick check and a proper one is the actual reason a scored report is worth paying for over a free five minute look.
The problem with checking on your own logged in account
We covered this in detail in our post on the ChatGPT personalisation problem, but the short version is that a logged in AI account carries memory, prior conversation history, and inferred preferences that quietly shape the answer you get. If you've ever mentioned your own business, your industry, or even your general location in a previous chat, that context can leak into how the model answers a fresh question, in either direction. A logged in test can make your business look stronger than it actually is to a genuine prospective customer, or in some cases weaker, and either way it isn't measuring what an anonymous customer would actually see.
Why one prompt isn't enough either
A single question, however well chosen, only tells you how a model responds to that exact phrasing on that exact day. Real customers ask the same underlying question a dozen different ways: by location, by specific need, by comparison against a named competitor, by occasion. A business can rank well on one phrasing and disappear entirely on a slightly different one that a real customer is just as likely to use. Testing a single prompt and generalising from it is a bit like judging a shop by what's visible through one window.
Why four platforms, not one
ChatGPT, Gemini, Claude and Perplexity are built differently, draw on different sources, and behave differently even when asked the same question. We've documented this directly: golf operators in our four operator study performed completely differently across just two of these platforms, with the same business absent on one and ranked highly on another. Testing only the platform you personally happen to use tells you nothing reliable about the other three, and a customer has no way of knowing which one a prospective customer will open.
What proper testing actually looks like
- Logged out sessions, every time, so results reflect what an anonymous customer sees, not a personalised or history influenced answer.
- A fixed panel of real customer style prompts, covering different phrasings, occasions and comparison angles within the same category, rather than one representative question.
- All four major platforms tested directly, plus Google AI Overviews, since each has a genuinely different mechanism and can produce a genuinely different answer.
- Results logged as evidence, not summarised into a number alone. Being able to see the actual conversation a model produced is what makes a fix list credible rather than generic.
Where to start
If you've only ever checked your own visibility on your own logged in ChatGPT account, that check was measuring something closer to your personal AI history than your actual market visibility. It's a reasonable first look, but not a reliable one, and it's worth treating any result from it as a starting question rather than an answer.