Case Study

Why we test AI logged out, across four platforms, every time

It's easy to check whether AI recommends your business. Open ChatGPT, type a question, read the answer. It takes thirty seconds and it feels definitive. It's also, on its own, one of the least reliable ways to test this, and it's worth explaining exactly why, because the gap between a quick check and a proper one is the actual reason a scored report is worth paying for over a free five minute look.

The problem with checking on your own logged in account

We covered this in detail in our post on the ChatGPT personalisation problem, but the short version is that a logged in AI account carries memory, prior conversation history, and inferred preferences that quietly shape the answer you get. If you've ever mentioned your own business, your industry, or even your general location in a previous chat, that context can leak into how the model answers a fresh question, in either direction. A logged in test can make your business look stronger than it actually is to a genuine prospective customer, or in some cases weaker, and either way it isn't measuring what an anonymous customer would actually see.

Why one prompt isn't enough either

A single question, however well chosen, only tells you how a model responds to that exact phrasing on that exact day. Real customers ask the same underlying question a dozen different ways: by location, by specific need, by comparison against a named competitor, by occasion. A business can rank well on one phrasing and disappear entirely on a slightly different one that a real customer is just as likely to use. Testing a single prompt and generalising from it is a bit like judging a shop by what's visible through one window.

Why four platforms, not one

ChatGPT, Gemini, Claude and Perplexity are built differently, draw on different sources, and behave differently even when asked the same question. We've documented this directly: golf operators in our four operator study performed completely differently across just two of these platforms, with the same business absent on one and ranked highly on another. Testing only the platform you personally happen to use tells you nothing reliable about the other three, and a customer has no way of knowing which one a prospective customer will open.

This is the actual difference between a free checker and a proper diagnostic. It's not that free tools are dishonest, it's that a single logged in query on a single platform is structurally unable to answer the question a business actually needs answered: what does a real, anonymous customer see, on whichever AI they happen to be using, across the range of ways they're likely to ask.

What proper testing actually looks like

Where to start

If you've only ever checked your own visibility on your own logged in ChatGPT account, that check was measuring something closer to your personal AI history than your actual market visibility. It's a reasonable first look, but not a reliable one, and it's worth treating any result from it as a starting question rather than an answer.

Get a properly controlled test, not a personalised guess.

Our AI Recommendation Score is built on logged out, multi platform, multi prompt testing, with the actual conversations included as evidence. Start with a free AI Snapshot.

Get my free AI Snapshot