Code & data · robots.txtrobots.txt generator for AI bots and Yandex

Build a robots.txt that lets search engines in, decides on AI bots class by class and talks to Yandex correctly. Paste a ready file — we will check who it actually lets in.

The file is built in your browser — neither rules nor addresses go to a server
01Generator02About robots.txt03Questions04Related
01

Generator

Basic rules

Clean-param for Yandex · parameters · path
Presets

Model training

GPTBot, ClaudeBot, Google-Extended, CCBot…

Blocking affects neither search nor answers with links.

GPTBot
OpenAI
follows robots.txtdocumentation
ClaudeBot
Anthropic
follows robots.txtdocumentation
Google-Extended · token
Google
follows robots.txtdocumentation
Applebot-Extended · token
Apple
follows robots.txtdocumentation
CCBot
Common Crawl
follows robots.txtdocumentation
Meta-ExternalAgent
Meta
follows robots.txtdocumentation
Amazonbot
Amazon
follows robots.txtdocumentation
MistralAI-Training
Mistral
follows robots.txtdocumentation
Bytespider
ByteDance
no documentation
cohere-training-data-crawler
Cohere
no documentation
Timpibot
Timpi
no documentation

AI search and citation

OAI-SearchBot, Claude-SearchBot, PerplexityBot, YandexAdditional…

Blocking removes the site from answers with links — that is lost traffic.

User requests

ChatGPT-User, Claude-User, Perplexity-User…

Many of them may ignore robots.txt; only the server is reliable.

robots.txt · 27 lines
User-agent: *
Disallow: /admin/
Disallow: /search
# AI: model training — blocked
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: CCBot
User-agent: Meta-ExternalAgent
User-agent: Amazonbot
User-agent: MistralAI-Training
User-agent: Bytespider
User-agent: cohere-training-data-crawler
User-agent: Timpibot
Disallow: /
# Yandex: Neuro — allowed except the general blocks (it does not read *)
User-agent: YandexAdditional
User-agent: YandexAdditionalBot
Disallow: /admin/
Disallow: /search
Clean-param: utm_source&utm_medium&utm_campaign /
Sitemap: https://example.ru/sitemap.xml

What robots.txt cannot do

Google AI Overviews and Bing chat run on the ordinary search robots — limit them with meta tags in <head>.

Google AI Overviews and AI Mode
<meta name="robots" content="nosnippet"> or max-snippet:0
Bing chat answers
<meta name="robots" content="noarchive, nocache">
02

About robots.txt

robots.txt tells robots which parts of a site they may crawl. With the arrival of AI it got a new job: deciding whether a company may train a model on your texts and show them in chatbot answers. AI bots come in three kinds, and blocking each has a different effect — the generator shows it before you save the file.

Training

GPTBot, ClaudeBot
Disallow: /
Blocking stops the site's texts from being used to train models and affects neither search nor citation.

AI search

OAI-SearchBot, PerplexityBot
traffic
These bots bring the site into answers with links. Blocking removes you from such answers.

User requests

ChatGPT-User
may ignore
Bots that open a page at a person's request. Many state outright that they may not follow robots.txt.

Google-Extended

A token, not a robot
no effect on Search
Blocks Gemini training and grounding in Gemini and Vertex AI, but does not affect Google Search or AI Overviews — those run on the ordinary Googlebot.

YandexAdditional

Yandex
by name only
Stops the site from appearing in Neuro and Alice answers. Works only when named explicitly: it does not read rules for “*”.

No Host

Obsolete
Host:
Yandex dropped the Host directive long ago — the main mirror is set with a redirect. The generator does not add it.
03

Frequently asked questions

Block the “training” class in robots.txt: GPTBot, ClaudeBot, Google-Extended, CCBot, Meta-ExternalAgent and others. The generator puts them in one group with “Disallow: /” — that is the “Don't train, but cite” preset. Such a block affects neither search nor answers that link to your site: other bots handle those, and it leaves them alone.

Updated

The bot list is checked against company documentation once a month. robots.txt is a request, not protection: only the server closes a site reliably.

«» added to favorites