Citation volatility is the rate at which the set of cited sources changes across repeated runs of the same prompt on the same AI engine under the same conditions. It is a metric defined by blimpp, published here with its full measurement method. High volatility means the engine is drawing from a wide, unstable pool of sources for that question; low volatility means a settled source set that is harder to enter and harder to lose.
In one sentence
Citation volatility tells you whether the sources an engine cites for a question are stable or churning, which is the context every single-run screenshot is missing.
How citation volatility is calculated
Citation volatility = (1 − mean pairwise overlap of cited source sets across runs) × 100
For each pair of runs, overlap is the share of cited sources the two responses have in common relative to all sources either cites. Averaging across all pairs and subtracting from one gives the volatility percentage: 0% means every run cites an identical source set, 100% means no two runs share a source. Run counts and controls follow the repeated-run protocol.
Worked example
A prompt is run five times on one engine. Comparing the ten possible pairs of responses, the average overlap between cited source sets is 60%. Citation volatility for that prompt is 40%: on average, four in ten sources change between any two answers. Illustrative example: blimpp benchmark data is added to this page as studies publish.
Why citation volatility matters
Volatility is the honesty check on every AI visibility claim. Where volatility is high, a competitor’s presence in one answer is weak evidence of anything, and your absence from one answer is equally weak. It also shapes strategy: a high-volatility question is winnable quickly because the engine has not settled on its sources, while a low-volatility question means displacing an incumbent source set. Reporting volatility alongside probability is what separates measurement from anecdote.
What affects citation volatility
Evidenced factors: generative answering is a stochastic, multi-stage pipeline rather than a fixed ranking, which is why repeated identical prompts produce different source sets at all; this is documented in the research literature on generative visibility. Factors with practitioner evidence: query specificity, since broad commercial questions fan out into more sub-queries and more candidate sources; the depth of the available source pool for the topic; and engine-side model or index updates, which can reset a settled source set overnight. These are treated as working hypotheses until tested.
Related concepts
References
- Martinez, O. (2026), Optimizing Visibility in Generative Engines: A Critical Survey of GEO: arxiv.org
- Zhang, K. et al. (2026), From Citation Selection to Citation Absorption: arxiv.org
Author: Harpal Singh · Last reviewed: 7 August 2026