robots.txt and bot rules: asking versus enforcing
Two layers decide which AI visitors get in. Most businesses only look at one.
Am I Ready for Agents? editors · 29 September 2026 · Reviewed September 2026
robots.txt asks crawlers to stay out of some pages; bot rules at your CDN or host decide who actually gets through. The standard says robots.txt is not a form of access authorization, so a clear position needs both layers to say the same thing.
What does robots.txt actually do?
robots.txt is a plain text file at the root of a website. It lists user agent names, such as GPTBot or ClaudeBot, and the parts of the site each may or may not crawl. Crawlers that follow the standard read it before they visit other pages.
The standard is RFC 9309, the Robots Exclusion Protocol, and it is explicit about the limit: the rules are not a form of access authorization. A well-behaved crawler follows them. Nothing in the file stops one that does not.
Which AI visitors read it?
Vercel's verified bots directory, last updated on 10 September 2026, describes OpenAI's GPTBot as respecting robots.txt directives to exclude sites from training data, and lists separate names for search crawlers and for fetchers that act on a user's question. It also notes an exception worth knowing: Meta-ExternalFetcher, because a user started the fetch, may bypass robots.txt rules.
So a robots.txt line is a reliable signal to the crawlers that honour it, and only a request to everyone else.
What do bot rules at the CDN or host add?
Bot management sits in front of your site and decides, request by request, whether to allow, challenge or block. Cloudflare's AI Crawl Control shows which AI crawlers visit, tracks whether they follow robots.txt and lets you allow or block them. Vercel offers an AI bots managed ruleset that can log or deny known AI bots.
These rules enforce. They are also where accidents happen: a setting switched on to reduce scraping can turn away the agents acting for your customers.
Why must the two layers agree?
Suppose a publisher's robots.txt welcomes AI search crawlers while a firewall rule blocks every AI bot. The written position says one thing and the site does another, and nobody notices because different teams set them at different times. The reverse also happens: robots.txt refuses training crawlers, but nothing enforces that for the ones that ignore the file.
The self-check asks about both layers in the find area: whether someone decided what robots.txt allows (question 1), and whether anyone has checked what the CDN, firewall or bot protection blocks (question 2).
Is there a middle ground between allow and block?
Yes, and operators increasingly support it. Google says blocking Google-Extended does not affect inclusion or ranking in Google Search, and Cloudflare reports the same split for Applebot-Extended. In September 2026 Cloudflare added a Disallow AI Training setting that writes robots.txt rules to refuse training while keeping a site discoverable in search.
Verification adds another option: allow a named agent only when its identity is confirmed. Lesson 8 covers how that works.
What should you ask?
Two questions from the find group on ask your web team: what does our robots.txt say today about the main AI crawlers, and who decided it; and does our bot protection have an AI rule switched on, set to log, challenge or block. If the two answers do not match, you have found your first fix.