GLOSSARY · 42 TERMS · PLAIN ENGLISH

AI agent glossary: 42 terms in plain English

Short definitions for the words your web team will use.

Am I Ready for Agents? editors · Reviewed September 2026

The short of it

These are the terms most likely to come up when you discuss AI agents with developers or vendors. Each is defined in one or two plain sentences, with a link to the standard or source where one exists.

A2A, Agent2Agent protocol

A protocol that lets one organisation's agent find another agent and hand it a task. Since August 2026 it is part of the Agentic AI Foundation.

Source: A2A protocol

Accountable crawler

A Cloudflare designation, introduced in September 2026, for AI crawler operators that meet stated opt-out and reporting requirements.

Source: Cloudflare, mixed-use AI crawlers

Agent readiness

How easily AI agents can find your business, read what you offer, complete a task, pay, and be trusted by you, and you by them.

See: Self-check

Agentic AI Foundation

A foundation that the A2A project describes as directed by the Linux Foundation and home to MCP. A2A joined it in August 2026.

Source: A2A joins the Agentic AI Foundation

Agentic commerce

Buying and selling in which an AI agent does some or all of the shopping for a person.

See: Ecommerce

AI agent

Software that takes a goal from a person and carries out the steps on the web, such as searching, filling in forms or paying.

See: What agents do

AI Crawl Control

Cloudflare's feature for managing AI crawlers: it shows which ones visit, tracks whether they follow robots.txt and lets you allow or block them. Cloudflare says it is available on all plans.

Source: Cloudflare AI Crawl Control

AI crawler

A program run by an AI company that visits web pages automatically, either to train models or to build a search index.

Bot management

Tools at your CDN or host that identify automated visitors and decide which to allow, challenge or block.

Source: Vercel, Bot Management

Browser agent

An AI agent that works through a web browser, reading pages and using forms and buttons the way a person would.

See: Why AI agents get stuck

CAPTCHA

A challenge meant to tell people from automated visitors, such as picking images or ticking a box. It can also stop agents that are acting for real customers.

See: Forms, sign-in and consent

CDN

Content delivery network: a service in front of a website that serves pages quickly and filters traffic. Cloudflare is one; many hosts include one.

Content negotiation

A standard web mechanism in which a visitor says which format it prefers, for example markdown rather than HTML, and the server answers in that format if it can. Cloudflare's Markdown for Agents responds to requests that ask for text/markdown.

Source: Cloudflare, Markdown for Agents

Content signals

Short statements, often sent in a response header, saying whether content may be used for AI training, for search, or as input to AI answers.

Source: Cloudflare, Markdown for Agents

Google-Extended

A name Google provides so sites can opt out of helping improve Gemini Apps and Vertex AI without affecting Google Search.

Source: Vercel verified bots directory

Guest checkout

Buying or booking without creating an account first. For agents, it removes a sign-in step they may not be able to complete.

See: Ecommerce

Headless browser

A web browser run by software without a screen. Browser automation platforms such as Browserbase run them for their customers, including to submit forms.

Source: Vercel verified bots directory

HTTP Message Signatures

A standard, RFC 9421 (February 2024), for signing parts of a web request so the receiver can check who sent it. Web Bot Auth is built on it.

Source: RFC 9421

JavaScript-only content

Text that appears only after scripts run in the browser. Visitors that do not run scripts, including many automated readers, may see little or nothing.

JSON-LD

A common way to add structured data to a page as a block of labelled facts that machines read directly, often using Schema.org vocabulary.

See: Structured data

llms.txt

A proposed short guide at the root of a website that tells language models what the site is and which pages matter. Version 2 was published in August 2026.

Source: llmstxt.org

MCP server portal

A single approved entry point for the MCP servers an organisation uses, with logging of tool activity. Cloudflare made its portals generally available in September 2026.

Source: Cloudflare changelog, MCP server portals

MCP, Model Context Protocol

An open standard for connecting AI applications to tools and data. A business can run an MCP server so agents can use its services directly.

Source: MCP specification

OAuth

A standard way to let an app or agent act on an account after the owner approves it, without handing over a password.

Optional OAuth scopes

Permissions on a consent screen that a person can untick, so an app or agent gets only the access a task needs. Cloudflare introduced them in August 2026.

Source: Cloudflare, task-based OAuth consent

Pay per crawl

A Cloudflare feature, in beta, that lets site owners charge AI crawlers for access to their content.

Source: Cloudflare AI Crawl Control

Readiness scanner

A tool that visits your public pages and reports which agent-related checks pass. Scanners work on public pages, not pages behind a login.

See: Tools that check

Reverse DNS verification

Checking that a bot's IP address resolves back to a domain its operator controls. Vercel lists it as one of three ways platforms verify bots.

Source: Vercel, Bot Management

robots.txt

A file at the root of a website that asks crawlers to stay away from some pages. It is a request, not a lock.

Source: RFC 9309

Search crawler, AI

An AI crawler that builds the index behind an AI search product, such as OAI-SearchBot for ChatGPT search.

See: What agents do

Self-assessment

A score of your own answers about your business, like the self-check on this site. It measures what you know, not what your site does.

See: Readiness levels

Server-side rendering

Building a page's full HTML before it is sent, so every visitor, including agents that do not run scripts, receives the text.

Sitemap

A file, usually sitemap.xml, that lists the pages a site wants automated visitors to find.

See: Self-check question 3

Structured data

Labels in a page's code that state facts such as price, availability or date in a form machines read directly, often using Schema.org vocabulary.

Training crawler

An AI crawler that collects pages to train or improve AI models, such as GPTBot or CCBot.

See: Which AI bots to let in

User agent

The name an automated visitor gives when it requests a page, such as GPTBot or ClaudeBot. Any script can copy a name, which is why verification matters.

See: What agents do

User-initiated fetcher

An AI visitor that reads a page because a person asked a question, rather than crawling on a schedule, such as ChatGPT-User or Perplexity-User.

Verified bot

An automated visitor whose identity a CDN or host has confirmed, for example by its IP addresses or a cryptographic signature.

Source: Vercel, verified bots

Weakest-area rule

The rule in this site's rubric that your level can be at most one above the level your lowest area would get on its own.

See: Readiness levels

Web Bot Auth

A method for bots to sign each request so a site can check who operates them, built on HTTP Message Signatures (RFC 9421).

Source: Cloudflare, Web Bot Auth

WebMCP

A proposal that lets a web page offer tools, such as search or book, directly to an AI agent working in the visitor's browser.

Source: WebMCP proposal

x402

An open standard that uses the web's 402 Payment Required code so software can pay for something and then fetch it.

Source: x402.org