What AI agents do on the web today
Who is knocking, what each visitor wants, and which ones you can turn away without turning away the rest.
Am I Ready for Agents? editors · Reviewed September 2026
Four kinds of AI visitor reach business websites: crawlers that collect pages to train models, crawlers that build AI search indexes, fetchers that read a page when a person asks a question, and agents that act on a page for a person. A few companies run most of them, but each has its own name, and most can be allowed or refused separately.
Which AI visitors reach a website?
Public bot directories group AI visitors by the job they do. The table below uses the descriptions in Vercel's verified bots directory, which was last updated on 10 September 2026. Names in the second column are the user agent names your web team sees in logs and uses in robots.txt.
| Kind | Examples (operator) | What the directory says |
|---|---|---|
| Training crawlers | GPTBot (OpenAI), ClaudeBot (Anthropic), CCBot (Common Crawl), Meta-ExternalAgent (Meta) | GPTBot crawls web content to improve OpenAI's generative AI models and ChatGPT and respects robots.txt directives to exclude sites from training data. CCBot crawls for AI training and research. Meta-ExternalAgent crawls for uses such as training AI models or improving products. |
| AI search crawlers | OAI-SearchBot (OpenAI), PerplexityBot (Perplexity), Claude-SearchBot (Anthropic), Meta-WebIndexer (Meta) | OAI-SearchBot indexes websites for inclusion in ChatGPT's search results and does not crawl content for AI model training. PerplexityBot does the same for Perplexity's search results. Claude-SearchBot works to improve search result quality for users. |
| Fetchers acting on a user's question | ChatGPT-User (OpenAI), Claude-User (Anthropic), Perplexity-User (Perplexity), DuckAssistBot (DuckDuckGo), Meta-ExternalFetcher (Meta) | These fetch a page when a person asks something. The directory describes ChatGPT-User and Perplexity-User as not used for automated crawling or AI training. It notes that Meta-ExternalFetcher, because a user started the fetch, may bypass robots.txt rules. |
| Agents that act for a person | ChatGPT operator, listed as chatgpt-operator (OpenAI); Google-Agent (Google); ryebot (Rye) | Google-Agent navigates the web and performs actions upon user request, used by agents hosted on Google infrastructure such as Project Mariner. The chatgpt-operator entry handles user-initiated requests from ChatGPT operator and is not used for automated crawling or AI training. ryebot powers automated checkout on behalf of shoppers with explicit consent. |
Source: Vercel, Bot Management: verified bots directory · Read September 2026
What is the difference between a crawler and an agent?
A crawler collects pages in bulk, on its own schedule, to train a model or build an index. An agent works on one task for one person, at that person's request: it reads a few pages, fills in a form, compares options and, increasingly, tries to finish the job. The practical difference for a business is that blocking a crawler changes what an AI system knows about you later, while blocking an agent stops a customer's request now.
Can I refuse one kind and allow another?
Often, yes. Google describes Google-Extended as a separate product token that controls whether a site helps improve Gemini Apps and Vertex AI generative APIs, and says it does not affect a site's inclusion or ranking in Google Search. Cloudflare reports that Apple offers the same split with Applebot-Extended, and in September 2026 Cloudflare added a setting that blocks AI training through robots.txt while keeping a site discoverable in search.
robots.txt is a request, not a barrier: the standard, RFC 9309, says its rules are not a form of access authorization. Enforcing a choice needs bot management at your CDN or host.
Sources: Vercel verified bots directory (Google-Extended entry) · Cloudflare, mixed-use AI crawlers · RFC 9309 · Read September 2026
How do agents prove who they are?
By name, and increasingly by signature. Vercel lists three ways platforms verify a bot: checking that requests come from the operator's published IP ranges, a reverse DNS lookup, and cryptographic verification with Web Bot Auth, which signs requests using HTTP Message Signatures (RFC 9421). Cloudflare now checks Web Bot Auth signatures automatically when operators submit bots to its directory. Verification is what lets a site welcome a known agent without opening the door to any script that borrows its name.
Sources: Vercel, Bot Management · Cloudflare, BotBase for Operators · RFC 9421 · Read September 2026
What are agents looking for when they arrive?
Cloudflare frames the agent web as readable, discoverable, callable and payable. Cloudflare's own scanner, Is It Agent Ready, groups its checks into discoverability, content accessibility, bot access control, protocol discovery and commerce. ora.ai asks four questions: can an agent find and recommend you, access your data and understand you, use you (operate, authenticate, integrate), and pay you. The wording differs, but the ground is the same, and it is the ground the self-check covers in plain language.
Sources: Cloudflare, The agentic internet · Is It Agent Ready · ora.ai methodology · Read September 2026
Can I see how many AI visitors I get?
Not from this site, and we do not publish estimates. Your CDN or host may show it: Cloudflare's AI Crawl Control analyzes AI crawler traffic and tracks robots.txt compliance, and Vercel's bot management lets teams identify verified bots by name and category. Ask your web team what your provider already reports.
Sources: Cloudflare AI Crawl Control · Vercel, verified bots · Read September 2026