Method

Why a single ChatGPT answer proves nothing

AI assistants never answer the same way twice. How to measure a brand’s visibility with repeated runs and confidence intervals.

By Sébastien Monnier · · 6 min read · Lire en français

Ask ChatGPT “who is the best plumber in Lille?”. Write down the businesses it names. Ask exactly the same question again. There is a good chance the list will change.

0%25%50%75%100%3 answers: between 9% and 91%30 answers: between 31% and 69%300 answers: between 44% and 56%
Same observed rate (50%), three levels of certainty: the more answers, the narrower the 95% interval (Wilson).

This is not a bug: language models generate their answers probabilistically. For a measurement tool, this has a simple and often ignored consequence: a single answer proves nothing.

The single-answer score trap

Imagine a tool that asks the question once and displays: “Your client is cited ✅”. The next week: “Your client is no longer cited ❌”. Has the client lost visibility? Maybe. Or maybe the two measurements are simply two draws of the same random process.

With a single answer, you cannot tell a trend from a fluctuation. And a client report that swings between alerts and good news at random ends up unread.

The solution: several runs and a margin of error

You have to think like a poll:

  1. Ask each question several times to each engine, over the same period.
  2. Count in how many answers the business is mentioned.
  3. Calculate a confidence interval around the observed rate.

A 95% confidence interval gives the range in which the true probability of being cited very probably lies. For proportions computed on small samples, the recommended method is the Wilson interval: unlike the “rate ± 2 standard deviations” approximation learnt at school, it always stays between 0 and 100% and behaves well when the rate is close to 0 or 100%.

What it looks like in practice

Mentioned in…Observed rate95% Wilson interval
1 answer out of 333%6% – 79%
5 answers out of 1533%15% – 58%
20 answers out of 6033%23% – 46%

The same 33% rate is worth something very different depending on the number of answers. With 3 answers, the true value could be anywhere between “almost never” and “almost always”. With 60, you know what you are talking about.

That is why Recoia’s figures are calculated over several questions × several runs × several engines, and always shown with their interval.

Alerts that make sense

The interval also helps decide when to raise an alert. A simple rule: a drop is only flagged if the new measurement falls clearly outside the previous one’s interval. Normal fluctuations trigger nothing; real breaks (a model update, a new competitor, a deleted page) stand out immediately.

What about the cost?

Asking each question three times triples the bill… unless you pool. The same normalised question, for the same engine, language and area, can be run once per period and serve every business concerned. That is the principle behind Recoia’s pooled panels: more rigour, at a lower cost per audit.

Key takeaways

  • An AI answer is a draw, not a measurement.
  • You need several runs and a confidence interval (Wilson) to draw a conclusion.
  • Showing the margin of error does not weaken a report: it makes it credible.

For the details of our method, see the Methodology page.

Be among the first to try Recoia

Businesses and agencies: we are opening in waves. People on the list get the launch price for 12 months.

You are

Double opt-in by e-mail. Unsubscribe in one click. Your data is never sold.