# Is your site visible to AI crawlers? The 5-minute audit

> A copy-pasted 2023 robots.txt block can erase you from ChatGPT and Perplexity. The 5-minute audit, the bots that matter, and the allow-search strategy.

- URL: https://www.agentskillpacks.com/blog/ai-crawler-access-5-minute-audit
- Author: İsmail Günaydın
- Published: 2026-06-10
- Updated: 2026-09-18

## Quick answer

Audit AI crawler access in three steps. Fetch yourdomain.com/robots.txt and check rules against the 12 AI bots that matter (GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, Google-Extended, CCBot, and friends). Decide separately for search bots versus training bots. Then give allowed crawlers a curated map via llms.txt. No robots.txt at all means everything is allowed, which for most sellers is the correct state.

In 2023, half the internet pasted the same "block all AI bots" robots.txt snippet during the great scraping panic. Reasonable at the time. But the bots in that snippet now decide whether your site appears in ChatGPT and Perplexity answers, and a growing slice of buyers reads those answers instead of clicking ten blue links. Plenty of site owners deleted themselves from that channel and never noticed.

The audit takes five minutes. Here is the whole thing.

## What does an AI crawler audit check?

An AI crawler audit checks whether the bots that feed AI answers are allowed to read your site, and whether each block was a decision or an accident. It has three parts. First, read your robots.txt and see which rules apply to each AI user agent. Second, split those bots into two groups: training bots that collect text for future models, and search bots that fetch pages to answer live questions with a citation. Third, decide each group on purpose. Blocking a training bot keeps your words out of future models and costs no traffic. Blocking a search bot removes you from AI answers, and from the referral clicks those answers send. Most sellers should allow the search bots. A site with no robots.txt at all is already fully open, which for most product and content sites is the correct state.

## How do I read my robots.txt for AI bots?

Open yourdomain.com/robots.txt in a browser. Three outcomes are possible.

It does not exist, a 404. That means every compliant crawler assumes full access. For most product and content sites this is correct, and your audit is nearly done.

It exists and never mentions an AI bot. Then AI crawlers fall under your wildcard rules, whatever applies to `User-agent: *`. Usually fine, worth confirming.

It exists and contains a wall of `User-agent: GPTBot, Disallow: /` blocks. This is where the 2023 snippet lives, and where the audit earns its five minutes, because the question is whether you still mean it.

One trap catches people who add bots on purpose. Per [RFC 9309](https://www.rfc-editor.org/rfc/rfc9309.html), a crawler that matches a named `User-agent` group ignores the `*` group completely. If you write `User-agent: GPTBot` / `Allow: /`, your `Disallow: /api/` from the wildcard group no longer applies to GPTBot. Repeat the private-path rules in every named group. We shipped this exact mistake and fixed it in September 2026.

Evaluating the rules by hand gets fiddly: group selection, longest-match precedence, Allow-beats-Disallow ties. The [AI crawler checker](/ai-crawler-checker) does the evaluation for any domain and shows the exact rule behind each verdict, so you can confirm every result against your own file.

## Which AI bots train models and which ones cite you?

The crawler list breaks into two camps, and the entire strategy lives in the difference. These are the 12 bots the checker tests, with each vendor's own documentation.

| User agent         | Company      | Job                       | If you block it                             |
| ------------------ | ------------ | ------------------------- | ------------------------------------------- |
| GPTBot             | OpenAI       | Training                  | Kept out of future models, still citable    |
| OAI-SearchBot      | OpenAI       | ChatGPT search index      | Dropped from ChatGPT search answers         |
| ChatGPT-User       | OpenAI       | Live fetch for a user     | ChatGPT cannot open your page on request    |
| ClaudeBot          | Anthropic    | Training                  | Kept out of future models                   |
| Claude-User        | Anthropic    | Live fetch for a user     | Claude cannot open your page on request     |
| PerplexityBot      | Perplexity   | Search index              | Dropped from Perplexity answers             |
| Perplexity-User    | Perplexity   | Live fetch for a user     | Perplexity cannot open your page on request |
| Google-Extended    | Google       | Gemini training control   | No effect on Search rankings                |
| Applebot-Extended  | Apple        | Apple AI training control | No effect on Siri or Spotlight search       |
| CCBot              | Common Crawl | Open web corpus           | Out of a corpus many models train on        |
| Bytespider         | ByteDance    | Training                  | Out of ByteDance models                     |
| meta-externalagent | Meta         | Training                  | Out of Meta models                          |

Sources: [OpenAI](https://developers.openai.com/api/docs/bots), [Anthropic](https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler), [Perplexity](https://docs.perplexity.ai/docs/resources/perplexity-crawlers), [Google](https://developers.google.com/crawling/docs/crawlers-fetchers/google-common-crawlers), [Apple](https://support.apple.com/en-us/119829), [Common Crawl](https://commoncrawl.org/ccbot). Anthropic also documents Claude-SearchBot, which indexes pages for Claude's search results. Treat it like OAI-SearchBot.

Blocking a training bot does not hide you from AI search. Blocking a search or fetch bot does, and that is the lever that costs traffic. The blanket 2023 snippet blocked both camps at once, which is exactly why it needs revisiting. The decision was never one decision.

## Should I block AI training bots?

Allow the search bots, decide separately about the training bots. That is the whole strategy for most sites.

For our shop the call was easy. We allow all twelve, because an AI assistant recommending our toolkits is free distribution, and our content earns by being found, not by being scarce. I documented the full reasoning, bot by bot, in [our robots.txt and llms.txt, explained](/blog/toolgenx-robots-and-llms-txt-explained).

The opposite call is just as rational for a different business. A newsroom or a paid-course shop selling the content itself can sensibly allow citation bots and block training bots, keeping the traffic channel while declining to feed the corpus. What is never rational is the third state most sites are actually in: a robots.txt that encodes a decision nobody remembers making.

## Do I need an llms.txt file?

Once crawlers can reach you, the next question is what they read first. An [llms.txt](https://llmstxt.org/) file at your site root is a curated markdown map: site name, one-sentence summary, and annotated links to the pages that answer "what is this site and can I trust it".

Honesty requires saying what it does not do. Google has stated Search ignores the file, and no platform treats it as a ranking input. What it does: AI coding assistants fetch it, some answer engines read it when summarizing a domain, and it is the one document where you control the exact wording an AI encounters first. Ten minutes of work for a maybe is a fair trade. The [llms.txt generator](/llms-txt-generator) builds a spec-valid file in the browser and validates it as you type.

## What does this audit not cover?

Access is step zero, not the strategy. Whether AI engines actually cite you once they can read you depends on citability: answer-shaped content blocks, quotable facts with numbers attached, schema markup, and entity signals. That is a content discipline, not a config file, and it is the gap where most "AI SEO" effort should actually go. The [AI Search Visibility Toolkit](/products/ai-search-visibility-toolkit) packages our full audit workflow for exactly that next layer: eleven skills, citability scoring included.

But run the five minutes first. There is no point optimizing for engines you blocked in 2023 and forgot about.

## FAQ

### Does blocking GPTBot remove my site from ChatGPT answers?

Not directly. GPTBot collects training data; ChatGPT's live answers with citations come through OAI-SearchBot and ChatGPT-User. Blocking GPTBot keeps you out of future training corpora while leaving you citable. Blocking the search bots is what removes you from answers, and traffic.

### Does blocking Google-Extended hurt Google rankings?

No. Google has documented that Google-Extended only controls Gemini model training. It is not a Search ranking signal, and AI Overviews use Googlebot. You can block Google-Extended and lose nothing in classic search results.

### My site has no robots.txt at all. Is that a problem?

No file means every compliant crawler assumes full access, which for most product and content sites is exactly the right state. It is the sites that copy-pasted a block-everything snippet during the 2023 panic that need the audit most, because they opted out of AI answers without deciding to.

### Do AI companies actually respect robots.txt?

The documented major bots do: OpenAI, Anthropic, Google, and Apple publish their user agents and honor the protocol. Reported violations cluster around smaller undisclosed scrapers. robots.txt is a published preference, not a wall; for hard enforcement you need CDN or WAF rules.
