Methodology · v3.0 · June 2026 edition
How the score is built.
One number, 0 to 100, answering a single question: does an AI assistant put this brand
in front of a buyer? It combines five measured dimensions, each capturing a distinct way
a brand can fail — or succeed — at being seen.
This page publishes the scoring function, the sampling design, the engines and the
limitations. No brand can pay to rank. Scores are estimates from a model, revised monthly,
and reported with confidence intervals.
312brands measured
15categories
4AI engines
2,160sampled responses / edition
01
The measurement problem
A large language model does not return the same answer twice. Ask it the same buyer question
three times and you may get three different sets of brands, in three different orders. This is
not a bug to be engineered away; it is how the systems work.
The consequence is unforgiving: a single query is not a measurement. Any
"AI visibility score" derived from one call to one model is a screenshot of a coin toss.
Most of what is currently sold as AI visibility measurement is exactly that.
Everything below follows from taking that problem seriously. We sample repeatedly, across
paraphrased prompts and across four independent engines, and we treat the spread of the
results as information rather than noise to be hidden. A brand that appears erratically is
not the same as a brand that appears reliably, even when their average is identical — and
the score says so.
02
Collection
Every edition is a fresh collection run. Nothing is carried over, adjusted or smoothed
against the previous month.
Sampling design
| Prompts per category (n) | 12 buyer-intent questions, held fixed across editions |
| Draws per prompt, per engine (k) | 3, in independent sessions |
| Engines (E) | 4 |
| Sampled responses per category (N) | n × k × E = 144 |
| Sampled responses per edition | 2,160 across 15 categories |
Prompts are held constant between editions so that score movement reflects a change in the
world, not a change in the question. Prompt-set revisions trigger a methodology version bump
and are logged below.
Engine manifest
Four assistants are queried, each through an independent lane running in parallel. Each lane
resolves the query with live web retrieval enabled, so that what we record is what a buyer
would actually see — not what the model recites from memory.
ChatGPTOpenAI · web-search enabled
PerplexitySonar · retrieval-native
GeminiGoogle · grounded generation
ClaudeAnthropic · web-search enabled
The exact model identifier and the access timestamp are recorded for every individual
response and retained with the run. This is what makes a score movement diagnosable after
the fact: when an engine changes underneath us, the log tells us when.
Pipeline
Collection is automated. A scheduled orchestration workflow fans the prompt set out across
the four engine lanes in parallel, captures each raw response, and passes it to an extraction
stage that resolves three things per response: which brands are named,
in what order, and which sources the engine cited in support.
Entity resolution maps surface forms to canonical brands, so that "Palo Alto", "Palo Alto Networks"
and "PANW" collapse to one record. Results are written to a versioned store; every score on this
site is traceable to the individual responses that produced it.
Exclusions, and why they exist
An automated extractor pulls every proper noun an engine emits. Left unfiltered, a question
about cybersecurity vendors returns AWS, Python and Gartner alongside the actual vendors, and
the ranking becomes meaningless. The pipeline therefore excludes, by rule:
- The generative engines themselves, and the model families behind them
- Generic cloud and hosting infrastructure, unless the category is cloud infrastructure
- Analyst and research firms cited as sources rather than as vendors
- Programming languages, protocols and open standards
Category membership is then reviewed by a human before publication. We consider this an
asset, not an embarrassment: an unsupervised taxonomy produces a clean-looking index that
is quietly wrong, and no serious benchmark ships one. The review adjusts who belongs in
a category. It never adjusts a score.
03
The scoring function
Five dimensions, each normalised to the interval [0,1], each weighted, summed and scaled to
0 to 100. The weights are fixed for the lifetime of a methodology version.
p Presence
Is the brand cited at all?
30% The share of sampled responses, across the full prompt set and all four engines, in which the brand is named. A brand absent from the answer scores zero here. Product quality is irrelevant if the model never surfaces it.
p = citations / N
N is the total number of sampled responses for the category. Bounded [0,1].
Where the brand lands within the answer when it does appear. Being named first is not the same as being buried eighth. Position is discounted logarithmically, the standard decay used in ranked-retrieval evaluation.
q = mean over citing responses of 1 / log2(1 + pos)
pos = 1 yields 1.00; pos = 3 yields 0.50; pos = 8 yields 0.32. Bounded (0,1].
e Engine coverage
On how many engines?
20% How many of the four engines cite the brand above a minimum reliability threshold. Broad coverage signals durable, structural visibility rather than a single-model artifact that vanishes at the next update. A single lucky hit does not count as coverage.
e = |{ engines citing the brand in >= 10% of their responses }| / 4 The threshold exists to separate structural visibility from sampling noise. Bounded [0,1].
c Consistency
Is it stable?
15% How stable citation is across paraphrased prompts and repeated draws. Because LLM outputs are non-deterministic, a brand that appears only intermittently has fragile visibility. We score that fragility explicitly rather than averaging it away.
c = 1 - (sigma / mu), clamped to [0,1]
sigma and mu are the standard deviation and mean of the per-run citation rate. This is the complement of the coefficient of variation.
a Source authority
Why is it cited?
15% The independence of the sources an engine leans on when it surfaces the brand. A citation grounded in editorially independent third-party coverage is a stronger signal than one grounded in the brand’s own marketing pages, and it is far more durable across model updates.
a = weighted mean of the independence tier of each cited source
Independent third-party editorial = 1.0 · industry directory or aggregator = 0.6 · brand-owned or self-published = 0.2.
Confidence
Presence is a binomial proportion measured over N sampled responses, so its uncertainty is
calculable rather than a matter of opinion. We compute the 95 percent Wilson score
interval on p and propagate it through the weighting to produce an interval on the
published score.
How to read the rankings. Where two brands' intervals overlap, they are tied.
They are not first and second. A three-point gap in the middle of a leaderboard is frequently
not a real difference, and we would rather say that than sell you a false decimal. Rank order
is meaningful at the top and bottom of a category; in the middle, read the interval, not the row.
04
What we publish, what we protect
A benchmark that hides its mathematics is asking to be trusted rather than checked.
The formula, the sampling design and the engine manifest are all above. What stays closed is
the material that would let the index be gamed rather than verified.
Published
- The complete scoring function and all five weights
- Sample sizes: prompts per category, draws per prompt, engines
- The engines measured, with model identifier and access date logged per run
- Confidence intervals on every score
- The exclusion rules applied to the extractor
- The sources behind each brand's citations
- Known limitations, in full
- The methodology version history, with every change logged
Protected
- The exact prompt strings, per category
- The training corpus behind the source-authority model
- The entity-resolution and anti-noise internals
The prompt strings stay closed for one reason: publishing them would let a brand optimise
against the twelve questions rather than against the category, and the index would stop
measuring anything. This is the same reason a held-out test set is never released.
05
Limitations
What this index cannot tell you. We publish this section because the alternative — implying
a precision we do not have — is how measurement gets a bad name.
We measure citation, not revenue.
A high score means AI assistants put a brand in front of buyers. It does not mean those buyers convert. Roughly 82 to 88 percent of AI citations produce no click, and no link inside an AI answer is taggable, so click-tracking captures under 20 percent of reality — for us and for everyone else. Anyone claiming a direct citation-to-revenue number today is guessing. Modelling that link is our v4.0 research objective, not a current capability.
The score is an estimate, with a confidence interval.
The Presence term is a binomial proportion over N sampled responses. We publish its 95 percent Wilson score interval alongside every score. Two brands whose intervals overlap should be read as tied, not ranked. Small gaps between adjacent brands are frequently not statistically meaningful, and we say so rather than manufacturing false precision.
Engines change under us.
Model versions, retrieval indexes and system prompts change without notice and without changelog. A score movement between editions can reflect a genuine change in a brand’s visibility, or a change in the engine. We log the model identifier and access date for every run so that shifts can be attributed after the fact — but we cannot always separate the two in real time, and we do not pretend otherwise.
The prompt set is finite.
We sample buyer-intent questions per category, not the full space of things a human might ask. A brand strong on queries outside our set will be under-measured. Widening prompt coverage is the single largest known source of improvement in the index.
Coverage is English and French, and B2B.
Source ecosystems differ by language and market. Scores are not portable across languages, and the index currently says nothing about consumer categories.
06
Conflict of interest
The LLM Visibility Index is built and operated by ZivRank, and ZivRank is measured
in it. It appears in the AI visibility agencies category alongside its direct
competitors. This is a conflict of interest and we state it plainly rather than leaving it
to be discovered.
Here is how it is contained:
- Same pipeline, no exception. ZivRank is scored by the same automated collection run, on the same prompt set, against the same four engines, with the same formula. There is no manual adjustment, and no separate path through the code.
- No pay-to-rank, including for ourselves. Nothing on this site can be purchased, and that constraint binds the operator first.
- Full decomposition, so you can check. Every brand's score, ZivRank's included, is published broken down into its five components with the citing sources attached. A reader who suspects the score is inflated can look at the underlying measurements and say so.
- Competitors are measured on their merits. Where a competitor scores above ZivRank, the index says so. The agency category is published in full, not filtered.
We considered excluding ZivRank from its own category. We rejected it: removing the data
point does not remove the interest, it only removes the reader's ability to audit it.
Disclosure plus reproducibility is the stronger safeguard, and it is the standard applied
by every credible index whose operator competes in a measured market.
07
The research programme
What the index measures today is what engines cite. That is a description of the past.
The open problems in this field are the two that follow — predicting citation before publication,
and connecting citation to revenue. Neither is solved, by us or by anyone. Here is what we are
building, and where it currently stands.
v3.1 In development
Learned source authority
Replace the heuristic independence tiers of dimension a with an authority score propagated over the citation graph itself — which sources co-occur, which sources are cited by other cited sources. Authority becomes measured rather than assigned. Target: a source-level authority model with published rank correlation against the current heuristic.
v3.2 In development
Citability model
A classifier that predicts P(cited) for a given page against a given buyer query, trained on the labelled citation record the index has been accumulating since 2024. Retrieval simulation plus a supervised head. Target metric: held-out AUC, published. This moves the index from describing the past to estimating the future — the open problem in this field, and the one no vendor has solved.
v4.0 Design stage
Incrementality
Causal estimation of the relationship between measured AI visibility and lagged business signals — branded search volume, demo requests, pipeline — using matched control groups and holdout experiments. Not click attribution, which is not technically possible inside an AI answer. Lift inference, with intervals. This is the hard problem the industry is currently papering over.
Each of these will ship with the same disclosure standard as the scoring function above: the
method published, the evaluation metric published, the failure modes published. A capability we
cannot evidence is a capability we will not claim.
08
Trust rules
- No pay-to-rank. Rankings derive only from observed engine outputs. No placement, tier or position is purchasable, by anyone, including the operator.
- Every score traces to its sources. Each brand's citations link to the material the engines relied on.
- Scores are estimates, and are reported with intervals. We do not publish a number without publishing its uncertainty.
- The methodology is versioned in public. Every change is logged below, with the edition it took effect in.
- ZivRank is measured on the same basis as every other company, and the conflict is disclosed above.
- Corrections are published, not silently patched. An error in the data is a news entry, not an overwrite.
09
Version history
v3.0Jul 2026
Full scoring function published. Sampling parameters, engine manifest and confidence intervals disclosed. Limitations and conflict-of-interest sections added. Research programme published.
v2.1Jun 2026
Engine coverage weighting refined. Vertical AI added as a tracked category.
v2.0Mar 2026
Source authority introduced as a fifth dimension. Prompt-variant sampling expanded.
v1.0Jan 2026
First public edition: presence, rank, engine coverage and consistency.
The LLM Visibility Index research programme has been running since 2024. The index was first
published in January 2026 and is operated by ZivRank SAS, incorporated in 2026.
10
Questions
Can a brand pay to rank higher in the LLM Visibility Index?
No. Scores are derived only from observed engine outputs. There is no paid placement, no sponsored tier, and no mechanism by which a brand can purchase a position. Brands are added to a category on the basis of category relevance, not commercial relationship.
How is the AI Visibility Score calculated?
The score is a weighted sum of five measured dimensions on a 0 to 100 scale: Presence (30 percent), Rank (20 percent), Engine coverage (20 percent), Consistency (15 percent) and Source authority (15 percent). Each dimension is normalised to the interval [0,1] before weighting. The full formula is published on this page.
Which AI engines does the index measure?
Four: ChatGPT, Perplexity, Gemini and Claude. Each is queried through its own lane in an automated collection pipeline, with the model identifier and access date logged for every run.
ZivRank operates the index and also appears in it. How is that handled?
ZivRank is scored by the same automated pipeline as every other brand, on the same prompt set, with no manual adjustment. Its category is disclosed on this page as a conflict of interest, and its underlying measurements are published in full so that any reader can check them. We consider disclosure plus reproducibility a stronger safeguard than exclusion, which would simply remove the data point without removing the interest.
How often are scores updated?
Monthly. Each edition is a fresh collection run, not an adjustment of the previous one. The methodology version in force is recorded against every edition.