Every AI visibility measurement is downstream of the prompts used to produce it, and a badly built prompt set can manufacture any finding you like. This page records how blimpp constructs prompt sets so that published results describe engine behaviour rather than prompt design. The companion rules for running those prompts live in the repeated-run protocol.
Intent classes
Prompt sets are built across three intent classes, reported separately: informational prompts asking how something works or what something means; commercial prompts asking what to buy, use or choose in a category; and comparative prompts asking engines to weigh named options. Class mix is stated with every finding, because engines cite different source types at different rates across the three, and a blended figure hides that structure.
Phrasing rules
Prompts are written the way real users ask, in natural language, not keyword strings. No prompt names the brand being measured unless the measurement is explicitly of branded queries, since naming a brand in the prompt contaminates every unbranded metric downstream. Leading phrasings that presuppose an answer are excluded. Where a prompt set derives from real query data, the derivation is stated; where it is constructed, the construction rules are stated.
Set sizes and stability
Category measurements use sets of at least 20 prompts; smaller exploratory sets are published as directional only, with their size stated. Prompts are never reworded mid-series: a changed prompt is a new prompt, tracked as such. Long-running fixed sets follow the rules in benchmark prompt sets.
Disclosure
Studies disclose either the full prompt set or, where publication would invite gaming of a tracked benchmark, the construction rules, intent mix and worked examples sufficient for another practitioner to build an equivalent set. Which of the two applies is stated in the study.
Related methodology
Author: Harpal Singh · Last reviewed: 7 August 2026