Most AI recommendation testing, including a lot of what's floating around online right now, gets run from whatever account the person testing happens to be logged into. That's a bigger methodological problem than it sounds, because ChatGPT with memory switched on doesn't answer the same question the same way twice, depending on what it already knows about the person asking. We ran a direct test to find out exactly how much that matters, and the result was not what we expected going in.
What we actually tested
During a recent panel test, the first pass on a corporate and work social event prompt was run from a personal, logged in account with an established history in the golf and business space. ChatGPT's answer referenced that background unprompted, and the venue being tested for a corporate event query was completely absent from the recommendation. Suspecting the account's memory was shaping the result, we re-ran the identical prompts logged out, with no history and no personalisation at all.
What held steady, and what changed
Not every prompt moved. The branded and location specific query, essentially "is this venue any good," returned an almost identical answer both logged in and logged out: same rough ranking, near identical wording, no material bias detected. That part of the test held.
The generic, buyer intent query was a different story. Logged in, the venue was absent entirely from the corporate event answer. Logged out, it appeared, correctly described, but never made it into the model's final practical recommendation once the answer moved from general description into a specific per group size suggestion. So the direction of the bias ran opposite to the obvious assumption: the personalised account suppressed the result on the open ended query, rather than inflating it, and the effect was concentrated entirely on the broader, less specific prompt rather than the narrow, branded one.
Why this happens
A model with memory of a user's background and interests appears to lean toward what it infers that specific person wants to hear, which on a broad, open ended query can crowd out businesses that would otherwise have surfaced on a genuinely neutral pass. On a narrow, branded query, there's less room for that inference to operate, because the question itself is already specific enough to anchor the answer regardless of who's asking.
Why this matters if you're testing your own visibility
If you check what AI says about your own business from your everyday, logged in account, you may be looking at a result shaped by your own search and chat history rather than a neutral read of what a genuine prospect would see. That cuts both ways: it could make your position look better than it really is, or in this case, considerably worse, depending on what that particular account happens to carry with it.
The practical takeaway isn't that personalisation is unbeatable. It's that broad, generic buyer intent queries, the ones with the most commercial value precisely because they're not brand searches, are also the ones most exposed to this kind of account level noise. The fix isn't a technical one. It's making sure your underlying signals, structured data, reviews, a clearly stated corporate or specialist offering, are strong enough that you don't need a favourable personalisation roll to get a fair hearing.
How we account for it now
Any prompt showing an unexpected result now gets a logged out re-test before it's treated as a finding, not just accepted from the first pass. It's a small extra step, and it's the difference between a real pattern and an artefact of one particular account's history.
Where to start
If you've ever checked your own AI visibility from a personal account and been reassured, or alarmed, it's worth a genuinely neutral re-check. What a real, first time prospect sees may look meaningfully different from what your own logged in session shows you.