BRIEFING · DECISION MEMO

Which AI bots should you let in? A decision memo for business leaders

Three positions you can take, what each one means, and the defaults you may already have without knowing it.

Am I Ready for Agents? editors · 17 September 2026 · Reviewed September 2026

The short of it

Treat AI bots as four kinds of visitor, not one. Most businesses that sell or take bookings will want agents acting for customers to get through. The real choice is about training crawlers, and several operators now let you refuse training without disappearing from their search.

Why is this not a single decision?

AI companies run several bots, and each has a different job. Vercel's verified bots directory, last updated on 10 September 2026, describes OpenAI's GPTBot as crawling to improve its generative AI models, OAI-SearchBot as indexing sites for ChatGPT's search results without crawling for training, and ChatGPT-User as handling requests a user starts. The directory lists similar splits for Anthropic and Perplexity.

A fourth kind acts for a person: the directory lists Google-Agent, which navigates the web and performs actions on a user's request, and ryebot, which powers automated checkout for shoppers with explicit consent. Blocking every AI bot blocks these too.

Position 1: allow everything

The simplest position, and the one many sites hold by default. Your content can be used for training, you appear in AI search, and agents acting for customers can read and use your pages. The cost is that you have no say in training use, and you depend on your bot protection to catch the scripts that only pretend to be well-known bots.

Position 2: refuse training, allow search and user requests

The middle position, and the one operators increasingly support. Google describes Google-Extended as a separate token that controls whether a site helps improve Gemini Apps and Vertex AI generative APIs, and says it does not affect inclusion or ranking in Google Search. Cloudflare reports the same split for Apple's Applebot-Extended.

In September 2026 Cloudflare added a Disallow AI Training setting that writes robots.txt rules to refuse training while keeping a site discoverable in search, together with an Accountable designation for crawler operators that meet stated opt-out and reporting requirements.

This position suits most publishers and many brands. It keeps you visible where customers look, and it keeps agents working.

Position 3: refuse all AI bots

Possible, but costly for anyone who sells or takes bookings. An agent acting for a customer cannot read your prices or complete your form, so that customer's task goes elsewhere. It is also harder to enforce than it looks: robots.txt is a request, and the directory notes that some user-initiated fetchers, such as Meta-ExternalFetcher, may bypass robots.txt rules because a person started the fetch.

What about agents that act for customers?

These are the visitors a business most wants to get right, and the hardest to identify. A user agent name can be copied by anyone. Verification closes the gap: Vercel lists IP range checks, reverse DNS and cryptographic signatures through Web Bot Auth, which signs requests with HTTP Message Signatures (RFC 9421). Cloudflare now validates Web Bot Auth signatures automatically when operators submit bots to its directory.

The practical position for stores and booking sites is to verify first, then decide. Let verified agents reach checkout, and apply your normal fraud rules to the purchase itself.

Which defaults might you already have?

Check three places. Your robots.txt, which may have been written years ago. Your CDN or host, which may have an AI bot rule switched on: Vercel's AI bots managed ruleset, for example, can log or deny known AI bots. And any content features you have enabled: Cloudflare's Markdown for Agents, when a site has not set its own policy, sends a content signal header saying yes to AI training, search and AI input.

None of these defaults is wrong. The point is to choose them on purpose.

How do you write the decision down?

  • Which kinds of AI visitor you allow, refuse or verify first.
  • The names your web team uses for each, from the operators' own documentation.
  • What your robots.txt says, and whether your bot protection enforces the same thing.
  • What content signals you send.
  • Who owns the policy and when it will be reviewed.

Bring the list to your web team with the find and trust questions.