Skip to content
AvocadoScore

Crawler directory · AI training crawler

GPTBot: user agent, robots.txt rules, and how to allow or block it

GPTBot collects public web pages that OpenAI may use to train its models. It is separate from OAI-SearchBot, which indexes pages for ChatGPT Search, so you can block one and allow the other.

Operator
OpenAI
Feeds
ChatGPT and OpenAI model training
Type
AI training crawler
robots.txt token
GPTBot

GPTBot user agent

Requests from GPTBot carry a User-Agent header like this (example — vendors change version numbers, so match on the token GPTBot, not the whole string):

Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.1; +https://openai.com/gptbot

To see how often it visits, filter your access log with grep -i "GPTBot" access.log. A User-Agent can be faked, so treat it as a claim, not proof.

Official documentation: platform.openai.com. Vendors change their crawlers, so check it for the current details.

Should you allow or block GPTBot?

Blocking stops future crawls from being used for model training. It does not remove content already collected, and it does not stop you being cited in live answers, which use separate retrieval bots.

robots.txt rules for GPTBot

Put these in the robots.txt at the root of your domain.

Block GPTBot

User-agent: GPTBot
Disallow: /

Explicitly allow GPTBot

User-agent: GPTBot
Allow: /

Block only part of your site

User-agent: GPTBot
Disallow: /private/
Allow: /

robots.txt states your policy; it does not enforce it. A CDN or firewall can block a bot your robots.txt allows, which is why it is worth testing rather than assuming.

Check what your site does today

The free AI bot checker reads your robots.txt and tells you whether GPTBot, ClaudeBot, PerplexityBot and nine other AI crawlers are allowed, blocked or crawling by default, then runs a live GPTBot fetch to prove your content is actually served.

Check my site free →

Questions, answered

What is GPTBot?+

GPTBot collects public web pages that OpenAI may use to train its models. It is separate from OAI-SearchBot, which indexes pages for ChatGPT Search, so you can block one and allow the other.

How do I block GPTBot in robots.txt?+

Add these two lines to the robots.txt at the root of your domain: User-agent: GPTBot Disallow: / User-agent tokens are matched case-insensitively. The rule applies from the crawler's next visit; it does not delete anything already collected.

Will blocking GPTBot affect my Google rankings?+

No. GPTBot is separate from Googlebot, so blocking it does not change how Google Search indexes or ranks you. Blocking stops future crawls from being used for model training. It does not remove content already collected, and it does not stop you being cited in live answers, which use separate retrieval bots.

What does the GPTBot user agent look like?+

An example User-Agent header is: Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.1; +https://openai.com/gptbot. Vendors change version numbers, so match on the token GPTBot rather than the whole string. A User-Agent can be spoofed, so verify important traffic against the vendor's published details before trusting it.

Related crawlers

More on the trade-offs in the AI bot policies guide and the robots.txt guide.