Claude is Anthropic’s AI assistant, which can search the web to ground answers in current sources and cite the pages it draws on. This is a platform reference page: how the surface works, how content becomes eligible to appear in it, and the controls publishers have, verified at the review date below.
In one sentence
Claude answers from training unless a question needs the live web, at which point its search retrieves and cites current pages, and a separate crawler decides whose pages are in that pool.
How Claude retrieves and cites
With web search engaged, Claude retrieves current pages relevant to the question, grounds its answer in them and links the sources used. Anthropic operates three distinct agents with independent robots.txt controls, mirroring the pattern OpenAI established. Claude-SearchBot indexes content to improve Claude’s search results: it is the visibility gate, the crawler that decides whether a site is citable inside Claude’s answers. ClaudeBot collects content that may be used for model training; blocking it is a training opt-out and does not remove a site from search. Claude-User fetches a page in real time when a user’s question requires it. All three honour robots.txt, including the crawl-delay directive.
What determines inclusion, and a housekeeping note
Eligibility for citation runs through Claude-SearchBot access, alongside the universal requirements: the site’s host, CDN and firewall must not block the crawlers above robots.txt, and pages must carry their substance in accessible rendered HTML. One housekeeping detail worth acting on: Anthropic’s older agent tokens, Claude-Web and anthropic-ai, are deprecated and no longer match any current crawler, so robots.txt rules targeting them do nothing and sites relying on them for control are not controlling anything. The training and search decisions are independent, so the common publisher stance of blocking training while staying citable is available here as elsewhere.
Why this surface matters
Claude carries a professional-heavy user base asking research and evaluation questions, and its clean separation of training, search and user-fetch agents makes it one of the more legible engines to manage access for. It also completes a five-engine measurement set: with per-engine behaviour differing and citation overlap often low, Claude’s answer set is a distinct territory to measure rather than an echo of the others.
Related concepts
References
- Anthropic’s crawler documentation, as reported by Search Engine Land, Anthropic clarifies how Claude bots crawl sites: searchengineland.com
Author: Harpal Singh · Last verified: 7 August 2026