The model says roughly the same thing every time. Your position here is real and you can plan against it.
One run is an anecdote. Twenty is a measurement.
Ask a model the same question twice and you can get two different answers. N-sampling runs your prompts repeatedly, then reports how much the answers actually varied, so the visibility number you take to a meeting is one you can defend.
Models are probabilistic, dashboards are not
Most AI visibility tools run a prompt once and present the result as a fact. That is a coin flip reported as a coin.
A language model samples from a distribution rather than reading from a record. Two identical prompts can return different vendors, different citations and a different ranking for you. Run once and you learn what happened that time, which is not the same as what happens.
That matters in both directions. A single bad run can send a team rewriting pages that were fine. A single good run can convince everyone a problem is solved when it never was.
Same prompt, N times, then compare
Set a depth for the brand, override it for the prompts that matter most, and run.
- 1
Pick a depth
Five, ten or twenty samples, set as a brand default and overridable per audit. Depth is a cost-versus-confidence decision and it stays yours.
- 2
Run the samples
The same prompt executes repeatedly against the same model. Every response is kept in full, not just its score.
- 3
Read the spread
Consistency, variance level, the range of visibility across runs, sentiment distribution and the hallucination rate across samples.
Filter completed audits by model, by variance level or by search, so you can pull up every contradictory result across your prompt set in one view.
Four bands, one of which needs you
Scored out of ten and labelled, so the list sorts by how unstable your position is rather than by how good the best run looked.
Small differences in wording and ordering, same substance. Normal model behaviour, nothing to chase.
Answers differ enough to change the impression a buyer walks away with. Worth reading the samples.
The model gave materially conflicting answers about you. Its sources disagree and nothing has become authoritative yet.
Contradictory is usually a sourcing problem rather than a content problem. The model is not confused about your writing, it is weighing sources that disagree, and no single one has become authoritative yet.
The spread, not just the average
Consistency and variance
A consistency percentage and a variance score out of ten, with the band that goes with it. The headline answer to whether this result is stable.
Visibility range
Average, minimum and maximum visibility across the samples, with the standard deviation. A 40% average from runs of 38 and 42 is a different story from one built out of 5 and 75.
Sentiment distribution
How the samples split across positive, neutral and negative, so you can see whether the model is inconsistent about facts or about tone.
Hallucination rate
The share of samples containing a false claim about you. A hallucination in one run of twenty is a different problem from one in eighteen.
A number you can defend in a meeting
The first question anyone senior asks about an AI visibility figure is how you know it is real. Sampling is the answer. It turns a score into a score with a confidence interval, and it tells you which of your prompts are stable enough to report and which are still noise.
It also protects your roadmap. Chasing a bad run is one of the most expensive ways to spend a content quarter, and a variance band is what stops it.
N-sampling, answered
Running the same prompt against the same model several times instead of once, then comparing the answers. Because models are probabilistic, one run tells you what happened once. Five, ten or twenty runs tell you what usually happens, which is the only version worth reporting to anyone.
Language models sample from a distribution rather than looking up a record. Two runs of an identical prompt can name different vendors, cite different sources and rank you differently. That is normal model behaviour, not a bug, and it is exactly why a single-run visibility score can be misleading in either direction.
A percentage showing how closely the samples agreed with each other. High consistency means your position is stable and you can plan against it. Low consistency means the model has no settled view of you, so today's good result is not something to build a forecast on.
Consistent, minor, moderate and contradictory. Contradictory is the one to act on: the model gave materially conflicting answers about you across runs, which usually means its sources disagree with each other and nothing has become authoritative yet.
Five is enough to catch obvious instability, ten is the practical default, and twenty is worth it for prompts tied directly to revenue. You can set a default depth for the brand and override it per audit.
Average, minimum and maximum visibility with the standard deviation across runs, average and range of mentions, average position, the sentiment split across samples, and the hallucination rate: the share of samples where a false claim about you appeared.
A hallucination that appears in one sample out of twenty is a different problem from one that appears in eighteen. Sampling gives you that denominator, so you can tell a rare misfire apart from something the model reliably believes.
N-sampling audits are on the Professional and Agency plans.
Find out how stable your AI visibility really is
Start with a free visibility audit, then sample the prompts that matter most to see whether the result holds.
Get a free visibility audit