SECTOR BRIEF · PUBLISHERS

What AI agents mean for publishers: training, search and live answers

The decision is not whether to allow AI, but which AI visitors, for which purpose.

Am I Ready for Agents? editors · Reviewed September 2026

The short of it

For publishers the question is which AI crawlers to let in and for what purpose. Operators increasingly separate training, search and live answers, so a publisher can refuse training while staying available to search and to readers' own requests. In September 2026 Cloudflare added a setting that does exactly that with robots.txt rules.

Which AI crawlers read publishers' pages?

All four kinds described on what agents do: training crawlers such as GPTBot, ClaudeBot and CCBot; AI search crawlers such as OAI-SearchBot, PerplexityBot and Claude-SearchBot; fetchers that read a page when a reader asks a question, such as ChatGPT-User and Perplexity-User; and research tools such as Gemini Deep Research, which the Vercel directory describes as analysing web content for multi-step research.

Source: Vercel verified bots directory · Read September 2026

Can I refuse training but stay in search?

Increasingly, yes. Google says Google-Extended controls whether a site helps improve Gemini Apps and Vertex AI generative APIs and does not affect inclusion or ranking in Google Search. Cloudflare reports that Apple offers the same with Applebot-Extended. In September 2026 Cloudflare added a Disallow AI Training setting that uses robots.txt rules and keeps a site discoverable in search, and an Accountable designation for crawler operators that meet stated opt-out and reporting requirements.

Remember that robots.txt is a request. RFC 9309 says its rules are not a form of access authorization, so enforcement needs bot management.

Sources: Cloudflare, mixed-use AI crawlers · Vercel verified bots directory · RFC 9309 · Read September 2026

What are content signals, and what does my site send today?

Content signals are short declarations of how content may be used: for AI training, for search, and as input to AI answers. One detail matters for publishers: Cloudflare's Markdown for Agents feature, when a site has not set its own policy, adds the header Content-Signal: ai-train=yes, search=yes, ai-input=yes. If you switch it on, decide your own signals first.

Source: Cloudflare, Markdown for Agents · Read September 2026

Can I charge AI crawlers?

Cloudflare offers Pay per crawl, which lets site owners charge AI services for access, as a beta feature within AI Crawl Control. AI Crawl Control itself, which manages AI crawlers, analyses their traffic and tracks robots.txt compliance, is available on all Cloudflare plans.

Source: Cloudflare AI Crawl Control · Read September 2026

Should articles be available as markdown or through llms.txt?

Both are options, not obligations. Markdown versions are easier for agents to read than full web pages. The llms.txt proposal, revised to version 2 in August 2026, lets a page point to its markdown version and to the llms.txt file that covers it. Whether to offer them depends on the position you take on AI use of your content.

Sources: llms.txt version 2 changes · Cloudflare, Markdown for Agents · Read September 2026