AI bot and llms.txt checker

Enter your domain to see what your robots.txt allows AI crawlers to do and whether your site has a valid llms.txt file. You can also generate a sample llms.txt to start from.

Generate an example llms.txt
# Site

> …

## Pages

- [Site](https://example.com)

Want a free SEO report for your site?

Our team reviews your rankings, technical health and opportunities and prepares a report for you.

A free SEO report for your site

Leave your details; our team will prepare a report on your rankings, technical health and opportunities and get in touch.

What it checks

AI companies visit the web for different reasons and under different user-agent names. Some collect content for model training, some build an index for search answers, and some fetch a page on the spot because a user asked for it in a chat. The tool looks at two files.

1. /robots.txt: For each crawler below, it shows whether your site allows access overall:

TokenProviderPurpose in brief (per provider documentation)
GPTBotOpenAICrawls content that may be used to train generative models
OAI-SearchBotOpenAISurfaces websites in ChatGPT's search features
ChatGPT-UserOpenAIVisits a page when a ChatGPT user asks for something
ClaudeBotAnthropicCollects content that may contribute to model training
anthropic-aiAnthropicA token not listed in Anthropic's current documentation but still common in robots.txt files
PerplexityBotPerplexitySurfaces websites in Perplexity search results
Google-ExtendedGoogleControls use of content for Gemini training and grounding; doesn't affect Google Search
Applebot-Extended, CCBot, Bytespider and othersVariousOther AI-related tokens frequently seen in robots.txt

2. /llms.txt: Does the file exist, and does it follow the basic format proposed at llmstxt.org? Under that proposal the file is Markdown, and the only required section is an H1 with the name of the site or project. It's followed by a blockquote summary, optional paragraphs of detail and H2 sections containing lists of links.

If you don't have one, the tool can generate a sample llms.txt from your site title and key pages as a starting point.

How to use it

  1. Enter your domain (for example example.com).
  2. The table shows "allowed", "blocked" or "partly blocked" for each crawler. Partly blocked means some directories are closed off.
  3. In the llms.txt section, see whether the file was found and how it fared against the format checks.
  4. If you like, copy the sample llms.txt, edit it and upload it to your site's root.

Interpreting the results

Training and search are separate decisions. OpenAI's crawler documentation explains that you can disallow GPTBot while allowing OAI-SearchBot, so content isn't used for training but can still appear in ChatGPT search. Anthropic's help article likewise defines separate tokens for ClaudeBot, Claude-User and Claude-SearchBot. Blocking everything with one line may cost you visibility you actually want.

User-initiated fetches behave differently. OpenAI notes that robots.txt rules may not apply to ChatGPT-User, and Perplexity's crawler page says Perplexity-User generally ignores robots.txt, because a person requested the fetch.

Google-Extended doesn't change your search rankings. According to Google's list of common crawlers, it's a control token only: it has no separate crawler and doesn't affect inclusion in Google Search.

Watch for accidentally blocking Googlebot. Broad rules written to keep AI bots out sometimes end up closing the site to every crawler under User-agent: *. Check the general rules in the report too.

Limitations

  • The tool reads robots.txt rules; it can't verify whether bots actually obey them. Your server logs are the more reliable source for that.
  • Providers update token names and purposes from time to time, so base decisions on their current documentation.
  • The llms.txt check covers the basic format only. It doesn't judge accuracy of the content or which services read the file.
  • It doesn't measure your visibility in AI answers. That's the job of the upcoming feature described on the AI visibility tracking page.

With SEOLOK

Are you mentioned in AI answers? Measure it every week

SEOLOK asks the questions your customers might ask to ChatGPT, Gemini and Google AI Mode every week. It checks whether your name appears in the answer, whether your site is cited as a source and who is recommended, and reports the result with its likely range instead of one confident number.

See how it works

Frequently asked questions

If I block GPTBot, will I disappear from ChatGPT search?

According to OpenAI's documentation, GPTBot relates to content that may be used for training generative models, while OAI-SearchBot is used to surface sites in ChatGPT's search features. You can block GPTBot and still allow OAI-SearchBot. Check the provider's current documentation for the exact effect of each setting.

Does blocking Google-Extended affect my Google Search rankings?

According to Google, no. Google-Extended manages whether content may be used for things like training Gemini models and grounding answers. It doesn't affect a site's inclusion in Google Search and isn't used as a ranking signal.

Is llms.txt required? Does Google use it?

It isn't required. llms.txt is a proposed format to help language models understand a site. We can't verify which AI services read it, and we're not aware of any official statement that Google uses it for Search. It's a cheap step with an uncertain effect.

Can a bot I've blocked in robots.txt still visit my site?

robots.txt is a request, not access control. Some providers state that robots.txt rules may not apply to fetches a user initiates directly. For a hard block you need measures at server or firewall level.

Does this tool measure my visibility in AI search?

No. It only checks your access rules and your llms.txt. A feature to measure how your site appears in AI answers is being prepared and will show up in the dashboard automatically once it launches.

Sources

  1. OpenAI: Overview of OpenAI crawlers
  2. Anthropic: Does Anthropic crawl data from the web, and how can site owners block the crawler?
  3. Perplexity: Perplexity crawlers
  4. Google Search Central: Google's common crawlers (including Google-Extended)
  5. Google Search Central: Introduction to robots.txt
  6. llmstxt.org: The /llms.txt file

Rank tracking

Why Google Rankings Drop: A Step-by-Step Checklist

An ordered checklist for a Google ranking drop: rule out reporting errors, technical issues, updates, competitors, seasonality and spam before you act.

SEOLOK Editorial Team · · 7 min read