Ask the Same Question Twice: What LLM Volatility Does to Casino Recommendations
Repeat the same question and the brand list changes. Why one screenshot proves nothing and what honest measurement looks like.
Here is an experiment anyone can run. Ask ChatGPT for the best online casino in your market. Note the brands. Open a new chat and ask the identical question. The lists will overlap, but they will not match. Run it a few more times and brands will drop out, reappear, and swap positions between runs. Now imagine building strategy on the first list alone. That is, in practice, what most teams do when they audit their AI visibility with a few manual prompts.
Why this happens
Language models sample. Temperature, retrieval variance, load balancing between model versions, all of it adds randomness to which brands make it into the final sentence. A brand can sit first in one run of a question and near the bottom in the next, with nothing about the brand changing in between. This is not a bug we expect vendors to fix soon, and academic work on generative engine volatility points the same direction. The volatility itself is not the scandal. Measuring through it incorrectly is.
What a single answer is worth
One answer is one sample from a distribution. It can catch a brand on its best run or its worst, and you cannot tell which from the screenshot. If your one manual check caught the good run, you filed a happy report. If it caught the bad run, someone started an unnecessary fire drill. Both reports are equally confident and equally worthless, because neither knows where in the spread it landed.
What honest measurement looks like
You treat every answer as one sample, never as the truth. Scores in the FoeGlass index are averages over repeat runs, so a brand that appears consistently across repeats scores very differently from a brand that flickered into one answer and never returned. Absence gets counted only when we witnessed the engine answering the question and your brand not being there. And every average stays connected to its verbatims, so you can see the spread with your own eyes rather than trust our arithmetic.
The practical takeaway is short. Any AI visibility claim based on single answers, screenshots included, is weather reporting from one glance out the window. If a number is going to reach your board deck, it should be an average over repeats with the sample size attached. That is the difference between measurement and anecdote, and in this field the anecdotes are unusually convincing.