Why AI answers change from one day to the next
Ask the same question twice and get different businesses named. That instability is real, it is measurable, and it changes how the number should be reported.
- Answers are assembled at request time, so variation between runs is normal behaviour.
- Four movers: retrieval variance, source changes, model updates, personalisation.
- One screenshot proves one run. It does not establish a position.
- Score on a majority of runs, report per surface, keep the raw responses.
The answers are assembled, not stored
An assistant is not looking up a saved answer. It retrieves sources, then generates a response from what it retrieved. Change the retrieval slightly and the answer changes.
So variation between runs is the normal behaviour of the system rather than a fault in it.
Four things that move between runs
Retrieval variance. A slightly different source set comes back, and a different business gets named.
Source changes. A comparison page or directory entered or left the set. Nothing about you changed.
Model and product updates. Shipped without notice and without a changelog you can read.
Personalisation and location. Reduced by signing out and setting location explicitly, never eliminated.
Why this makes a single check worthless
A screenshot proves that one run, on one surface, at one moment, named you. It does not establish that you are named, and an agency showing you one screenshot is showing you the best of several attempts.
The same logic applies in reverse. One run that omits you does not establish absence.
How to measure something that moves
Run each question multiple times per surface and score on a majority rather than a single result. Report per surface, never blended. Keep the raw responses so a disputed score can be checked.
Where results split, report the split as instability rather than rounding it away. A business sitting at the boundary of the consideration set is in a different position from one that is absent, and it is often the fastest thing to move.
Cadence follows from this
We run the full protocol quarterly rather than monthly, because these surfaces move on their own between runs and a monthly figure produces noise an owner would reasonably read as progress.
That is a slower reporting rhythm than most of the category offers. It is the one the measurement supports.
“An agency showing you one screenshot is showing you the best of several attempts.”
