AI engines are stochastic: the same prompt, on the same engine, minutes apart, produces different answers with different sources. Any metric built on a single response is therefore an anecdote wearing a number. The repeated-run protocol is the set of controls blimpp applies so that published figures such as citation probability and citation volatility describe an engine’s behaviour rather than one roll of it. It is published in full so another practitioner can approximate any blimpp measurement.
Controls held constant within a measurement
Engine and surface, since the same vendor’s assistant and search modes behave differently. Prompt wording, character for character. Locale and language. Session state, using clean or logged-out sessions wherever the surface permits, since personalisation contaminates comparison. Collection window, with all runs for a measurement gathered inside a stated period, because engine-side model and index updates can shift behaviour overnight. Where a surface exposes its model version, it is recorded; where it does not, that is recorded too.
Run counts
Minimum five runs per prompt per engine for any published probability, with the count always stated. Volatility measures use more runs where feasible, since pairwise comparisons are what the metric is built from. Small samples are published as directional with their limits stated, never silently blended into larger claims. One run is never published as a finding; single responses appear only as illustrations, labelled as such.
What is recorded per response
Engine, date and time, locale, prompt text, run number, the full cited source list, brands mentioned, brands recommended, and the response text where terms permit retention. Classification of those outcomes follows measuring citations, mentions and recommendations. Failed runs, refusals and empty retrievals are logged and reported, not discarded, since exclusion quietly biases every rate that follows.
Reporting rules
Every percentage is published with its numerator, denominator, engines and dates. Probabilities are reported alongside volatility for the same prompt set, because a 40% citation probability means something different in a stable source set than in a churning one. Cross-engine comparisons state each engine separately before any aggregate. Findings carry a limitations note covering at minimum: collection window, sample size, and any conditions that could not be controlled.
Known limitations
Engines update models and indices without notice, so measurements are snapshots of a moving system and are dated accordingly. Some surfaces cannot be fully de-personalised, which is recorded where it applies. Repetition inside one window measures short-run stability, not seasonal drift; tracking across windows is what the benchmark cadence exists for.
Related methodology
Author: Harpal Singh · Last reviewed: 7 August 2026