robots.txt · AI botsHow to block AI bots in robots.txt

A list of AI bots with their purpose and links to company documentation: who trains models, who brings readers from AI search and who may ignore your block.

The file is built in your browser — neither rules nor addresses go to a server
01Generator02Bot list03About the list04Questions05Related
01

Generator

Basic rules

Clean-param for Yandex · parameters · path
Presets

User-request bots often ignore robots.txt: only server rules close the site to them reliably.

Model training

GPTBot, ClaudeBot, Google-Extended, CCBot…

Blocking affects neither search nor answers with links.

GPTBot
OpenAI
follows robots.txtdocumentation
ClaudeBot
Anthropic
follows robots.txtdocumentation
Google-Extended · token
Google
follows robots.txtdocumentation
Applebot-Extended · token
Apple
follows robots.txtdocumentation
CCBot
Common Crawl
follows robots.txtdocumentation
Meta-ExternalAgent
Meta
follows robots.txtdocumentation
Amazonbot
Amazon
follows robots.txtdocumentation
MistralAI-Training
Mistral
follows robots.txtdocumentation
Bytespider
ByteDance
no documentation
cohere-training-data-crawler
Cohere
no documentation
Timpibot
Timpi
no documentation

AI search and citation

OAI-SearchBot, Claude-SearchBot, PerplexityBot, YandexAdditional…

Blocking removes the site from answers with links — that is lost traffic.

User requests

ChatGPT-User, Claude-User, Perplexity-User…

Many of them may ignore robots.txt; only the server is reliable.

robots.txt · 46 lines
User-agent: *
Disallow: /admin/
Disallow: /search
# AI: model training — blocked
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: CCBot
User-agent: Meta-ExternalAgent
User-agent: Amazonbot
User-agent: MistralAI-Training
User-agent: Bytespider
User-agent: cohere-training-data-crawler
User-agent: Timpibot
Disallow: /
# AI: search and citation — blocked
User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
User-agent: DuckAssistBot
User-agent: MistralAI-Index
User-agent: Amzn-SearchBot
User-agent: meta-webindexer
Disallow: /
# AI: user requests — blocked
User-agent: ChatGPT-User
User-agent: Claude-User
User-agent: Perplexity-User
User-agent: Meta-ExternalFetcher
User-agent: Amzn-User
User-agent: MistralAI-User
User-agent: Google-Agent
Disallow: /
# Yandex: Neuro and Alice answers — blocked
User-agent: YandexAdditional
User-agent: YandexAdditionalBot
Disallow: /
Clean-param: utm_source&utm_medium&utm_campaign /
Sitemap: https://example.ru/sitemap.xml

What robots.txt cannot do

Google AI Overviews and Bing chat run on the ordinary search robots — limit them with meta tags in <head>.

Google AI Overviews and AI Mode
<meta name="robots" content="nosnippet"> or max-snippet:0
Bing chat answers
<meta name="robots" content="noarchive, nocache">
02

AI bot list

from official company documentation
CompanyUser agentClassFollowsCheckedDocumentation
OpenAIGPTBotTrainingfollows robots.txt10.10.2026Open
AnthropicClaudeBotTrainingfollows robots.txt10.10.2026Open
GoogleGoogle-Extended
a token, not a separate robot
Trainingfollows robots.txt10.10.2026Open
AppleApplebot-Extended
a token, not a separate robot
Trainingfollows robots.txt10.10.2026Open
Common CrawlCCBotTrainingfollows robots.txt10.10.2026Open
MetaMeta-ExternalAgentTrainingfollows robots.txt10.10.2026Open
AmazonAmazonbotTrainingfollows robots.txt10.10.2026Open
MistralMistralAI-TrainingTrainingfollows robots.txt10.10.2026Open
ByteDanceBytespiderTrainingno documentation10.10.2026none
Coherecohere-training-data-crawlerTrainingno documentation10.10.2026none
TimpiTimpibotTrainingno documentation10.10.2026none
OpenAIOAI-SearchBotAI searchfollows robots.txt10.10.2026Open
AnthropicClaude-SearchBotAI searchfollows robots.txt10.10.2026Open
PerplexityPerplexityBotAI searchfollows robots.txt10.10.2026Open
DuckDuckGoDuckAssistBotAI searchfollows robots.txt10.10.2026Open
MistralMistralAI-IndexAI searchno documentation10.10.2026Open
AmazonAmzn-SearchBotAI searchfollows robots.txt10.10.2026Open
Metameta-webindexerAI searchfollows robots.txt10.10.2026Open
YandexYandexAdditional
by name only, ignores “*”
AI searchfollows robots.txt10.10.2026Open
YandexYandexAdditionalBot
by name only, ignores “*”
AI searchfollows robots.txt10.10.2026Open
OpenAIChatGPT-UserUser requestsmay not follow10.10.2026Open
AnthropicClaude-UserUser requestsfollows robots.txt10.10.2026Open
PerplexityPerplexity-UserUser requestsmay not follow10.10.2026Open
MetaMeta-ExternalFetcherUser requestsmay not follow10.10.2026Open
AmazonAmzn-UserUser requestsmay not follow10.10.2026Open
MistralMistralAI-UserUser requestsno documentation10.10.2026Open
GoogleGoogle-AgentUser requestsmay not follow10.10.2026Open
03

About the list

Every AI company names its robots in its own way and publishes crawling rules. We put them in one table and noted for each what it does and how it treats robots.txt — according to official documentation. Where there is no documentation, the table says so; every entry carries the date it was checked.

OpenAI

Three bots
GPTBot
GPTBot trains models, OAI-SearchBot brings the site into ChatGPT search, ChatGPT-User opens pages at a user's request and, according to OpenAI's documentation, may not follow robots.txt.

Anthropic

Three bots
ClaudeBot
ClaudeBot trains models, Claude-SearchBot handles search, Claude-User handles user requests. According to the docs, all three follow robots.txt.

Google

A token
Google-Extended
There is no separate AI robot: Google-Extended is a flag for training and grounding in Gemini and Vertex AI. AI Overviews are controlled by the nosnippet and max-snippet meta tags.

Perplexity

Search
PerplexityBot
The company's search bot. In 2025 Cloudflare reported that Perplexity got around blocks with undeclared crawlers.

Yandex

Neuro and Alice
YandexAdditional
Controls whether the site appears in generative answers. Works only when named explicitly.

Bing

Meta tags
noarchive
Microsoft has no separate AI bot. The nocache and noarchive meta tags limit use in Bing chat answers.
04

Frequently asked questions

The main ones: GPTBot, OAI-SearchBot and ChatGPT-User from OpenAI; ClaudeBot, Claude-SearchBot and Claude-User from Anthropic; Google-Extended from Google; PerplexityBot; Applebot-Extended; Meta-ExternalAgent; CCBot; YandexAdditional from Yandex. The full table is above: each bot has its company, class, attitude to robots.txt, a documentation link and the date it was checked.

Updated

The bot list is checked against company documentation once a month. robots.txt is a request, not protection: only the server closes a site reliably.

«» added to favorites