Perplexity is an AI answer engine that responds to queries with a synthesised answer built on live web retrieval, with inline citations to its sources. This is a platform reference page: how the surface works, how content becomes eligible to appear in it, and the controls publishers have, drawn from Perplexity’s own documentation and verified at the review date below.
In one sentence
Perplexity is retrieval-first by design: every answer is built from live sources it cites inline, which makes it the engine where fresh, citable pages convert to visibility fastest.
How Perplexity retrieves and cites
Perplexity operates its own search index, crawled by PerplexityBot, and composes answers from retrieved pages with numbered inline citations linking each claim to its source. A second agent, Perplexity-User, fetches pages when a user directly supplies or requests a URL; Perplexity treats this as user-initiated action rather than crawling. Perplexity states that it does not build foundation models, so allowing PerplexityBot is a search-inclusion decision, not a training decision; per its documentation, a site disallowing the crawler is not indexed in full, though the domain and a brief factual summary may still appear.
What determines inclusion
Eligibility runs through PerplexityBot access: robots.txt controls apply per agent, with changes reflected in up to about 24 hours, and Perplexity publishes WAF configuration guidance because firewall rules above robots.txt commonly block its crawlers silently. Publishers should know the compliance picture has been contested: a 2025 Cloudflare investigation reported undeclared Perplexity crawling activity, and Perplexity’s position is that user-initiated fetching is not bound by robots.txt. Practically, robots.txt governs the declared crawlers, and stricter control requires the server or WAF layer. Beyond access, Perplexity’s retrieval-first design rewards the general levers of citation probability, with recency weighted unusually heavily.
Why this surface matters
Perplexity is the most citation-dense of the major engines: sources are foregrounded in every answer rather than tucked behind it, so presence here is visible presence. Its freshness bias also makes it the fastest feedback loop for measuring whether new pages are entering AI answer sets at all, which is worth tracking per engine given how low cross-engine citation overlap can run.
Related concepts
References
- Perplexity, Perplexity Crawlers documentation: docs.perplexity.ai
- Perplexity, How does Perplexity follow robots.txt: perplexity.ai
Author: Harpal Singh · Last verified: 7 August 2026