Discoverability
Check whether AI search crawlers can access your docs
A repeatable check of robots.txt, CDN bot rules, status codes, and indexing signals for the crawlers behind ChatGPT search and Google's AI features.
By Hayden Bleasel10 min read

People reach your docs through a browser. AI search products reach them through a crawler, which can be turned away where a browser never is: a robots.txt group that names it, a CDN bot rule or challenge, a noindex header, or a canonical pointing elsewhere. Google's AI Overviews and AI Mode link only to pages eligible for ordinary Google Search, so they need Googlebot access and an indexed page. ChatGPT search needs OAI-SearchBot, a separate crawler with its own robots.txt rules.
This guide checks each layer against your live site and ends with a log of what you observed. It covers access and eligibility only: passing every check doesn't make any assistant show or cite a page.
Blume's defaults already let every crawler in: robots.txt allows all user agents, pages have self-referencing canonicals, and nothing is noindex unless you ask. So the block is usually in front of Blume, at your CDN or host, or in a robots.txt that replaced Blume's. The examples use docs.acme.example; swap in your domain.
Which crawler each AI search product uses
Each product runs separate crawlers for search and for training, and robots.txt controls each one:
| Where answers appear | Search crawler | Training crawler |
|---|---|---|
| Google AI Overviews and AI Mode | Googlebot | Google-Extended (a robots.txt token only) |
| ChatGPT search | OAI-SearchBot | GPTBot |
| Claude | Claude-SearchBot | ClaudeBot |
| Perplexity | PerplexityBot | None listed: Perplexity says PerplexityBot doesn't crawl for model training |
Google's AI features only need a page that's indexed and eligible to be shown with a snippet, and Google says Google-Extended doesn't affect inclusion in Search. OpenAI's crawler settings are independent: you can allow OAI-SearchBot and disallow GPTBot.
OpenAI, Anthropic, and Perplexity also run fetchers that visit a page when a user asks about it: ChatGPT-User, Claude-User, and Perplexity-User. OpenAI says ChatGPT-User isn't used to decide what appears in search. So when ChatGPT reads a link you paste, you've tested ChatGPT-User, not OAI-SearchBot.
Check the robots.txt crawlers get
Crawlers read the robots.txt your domain serves, which isn't always the one Blume built. Build the deployed commit and compare:
npx blume build
diff dist/robots.txt <(curl -s https://docs.acme.example/robots.txt)No output means they match. That path is for a static build; a server build writes the file to .vercel/output/static/ with vercel() and dist/client/ with cloudflare() or node(). If your host supplies the site URL at deploy time, a local build lacks the Sitemap: line; any other difference was added after the build, by your CDN or host.
Then test the live file against each crawler with robots-parser, the parser blume audit uses:
// Check which search and training crawlers the live robots.txt lets in.
// Usage: node check-robots.mjs https://docs.acme.example / /quickstart
import robotsParser from "robots-parser";
const crawlers = [
["Googlebot", "search: Google Search, AI Overviews, AI Mode"],
["OAI-SearchBot", "search: ChatGPT search"],
["Claude-SearchBot", "search: Claude"],
["PerplexityBot", "search: Perplexity"],
["GPTBot", "training: OpenAI"],
["ClaudeBot", "training: Anthropic"],
["Google-Extended", "training: Gemini"],
];
const [origin, ...paths] = process.argv.slice(2);
if (!origin) {
console.error("Usage: node check-robots.mjs <origin> [path ...]");
process.exit(1);
}
const robotsUrl = new URL("/robots.txt", origin).href;
const response = await fetch(robotsUrl);
console.log(`${robotsUrl}: HTTP ${response.status}`);
if (!response.ok) {
process.exit(1);
}
const robots = robotsParser(robotsUrl, await response.text());
for (const path of paths.length > 0 ? paths : ["/"]) {
console.log(`\n${path}`);
for (const [agent, purpose] of crawlers) {
const url = new URL(path, origin).href;
const verdict = robots.isAllowed(url, agent) ? "allowed" : "BLOCKED";
const line = robots.getMatchingLineNumber(url, agent);
const where = line > 0 ? ` (robots.txt line ${line})` : "";
console.log(` ${verdict} ${agent.padEnd(16)} ${purpose}${where}`);
}
}npm install robots-parser@3.0.1
node check-robots.mjs https://docs.acme.example / /quickstartHere's the output for a robots.txt meant to opt out of training that named OAI-SearchBot in GPTBot's group:
User-agent: *
Allow: /
User-agent: GPTBot
User-agent: OAI-SearchBot
Disallow: //quickstart
allowed Googlebot search: Google Search, AI Overviews, AI Mode (robots.txt line 2)
BLOCKED OAI-SearchBot search: ChatGPT search (robots.txt line 6)
allowed Claude-SearchBot search: Claude (robots.txt line 2)
allowed PerplexityBot search: Perplexity (robots.txt line 2)
BLOCKED GPTBot training: OpenAI (robots.txt line 6)
allowed ClaudeBot training: Anthropic (robots.txt line 2)
allowed Google-Extended training: Gemini (robots.txt line 2)A crawler follows only the group that names it most specifically, so once a group names OAI-SearchBot, the User-agent: * group stops applying to it.
To allow search and opt out of training, ship a public/robots.txt. It replaces Blume's file, so include the Sitemap: line:
User-agent: *
Content-Signal: search=yes, ai-input=yes, ai-train=no
Allow: /
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
Disallow: /
Sitemap: https://docs.acme.example/sitemap.xmlSet agents.contentSignals to { aiTrain: false } too, so agent-readability.json states the same policy. The Content-Signal line is a preference; the Disallow groups are what these crawlers' operators document.
Check status codes and bot challenges
Next, request a page as each crawler. Pass a plain sentence from it, with no links or formatting inside, so the script can confirm the text arrived:
#!/usr/bin/env bash
# Request one page as a browser and as three search crawlers, and print what
# each got back. Usage: ./check-access.sh <page URL> "<a plain sentence from the page>"
url="$1"
phrase="$2"
agents=(
"browser|Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36"
"Googlebot|Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)"
"OAI-SearchBot|Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36; compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot"
"PerplexityBot|Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot)"
)
for entry in "${agents[@]}"; do
name="${entry%%|*}"
headers=$(mktemp)
body=$(curl -s -A "${entry#*|}" -D "$headers" "$url")
status=$(head -n 1 "$headers" | tr -d '\r' | sed 's/ *$//')
if grep -qF "$phrase" <<<"$body"; then text="text found"; else text="TEXT MISSING"; fi
echo "$name: $status, $text"
grep -iE '^(server|cf-mitigated|x-robots-tag|retry-after|location):' "$headers" | tr -d '\r' | sed 's/^/ /'
rm -f "$headers"
doneThis site's own audit page passes on every row:
$ chmod +x check-access.sh
$ ./check-access.sh https://useblume.dev/docs/cli/audit "It crawls the HTML in"
browser: HTTP/2 200, text found
server: cloudflare
Googlebot: HTTP/2 200, text found
server: cloudflare
OAI-SearchBot: HTTP/2 200, text found
server: cloudflare
PerplexityBot: HTTP/2 200, text found
server: cloudflareA page behind a Cloudflare bot challenge returns this instead:
OAI-SearchBot: HTTP/2 403, TEXT MISSING
cf-mitigated: challenge
server: cloudflareCloudflare sets cf-mitigated: challenge on every challenge page. A 429 with Retry-After is a rate limit. A 200 with TEXT MISSING means the body isn't the page you expected. Blume's own rate limit covers the assistant, the playground proxy, and server-side search, never your pages.
Read the results as hints. Real crawlers are verified by IP address, and your request isn't. Vercel's Bot Protection, for example, challenges curl posing as a browser but lets verified bots through, while a rule that matches verified bots or their IP ranges never matches your request. So compare rows. If only one crawler's row fails, a rule on its user agent is the likely cause, and it matches the real crawler too. If every row fails, you've hit a general bot check that may still let verified crawlers through.
Then read the settings that decide it:
- Cloudflare. In Security Settings, Configure AI bot policies blocks or allows Search, Agent, and Training crawlers separately. Blocking Training also blocks crawlers classified as both search and training. AI Crawl Control allows or blocks each AI crawler.
- Vercel. In Firewall, the AI Bots Ruleset covers training, search, and user-triggered bots together, so Deny blocks OAI-SearchBot along with GPTBot. To stop only training, use robots.txt or a custom rule on their user agents.
- Elsewhere. Look for WAF rules on user agent, country, or ASN, and for site-wide bot challenges.
Check indexing and canonical signals
A fetchable page still stays out of results if it says noindex or its canonical points elsewhere. Check both, plus the header that overrides the HTML:
site=https://docs.acme.example
url="$site/quickstart"
curl -s "$url" | grep -oE '<link[^>]*rel="canonical"[^>]*>|<meta[^>]*name="robots"[^>]*>'
curl -sI "$url" | grep -i '^x-robots-tag' || echo "no X-Robots-Tag header"
curl -s "$site/sitemap.xml" | grep -c "<loc>$url</loc>"For an indexable page, such as this site's audit page, that prints a self-referencing canonical, no header, and one sitemap entry:
<link rel="canonical" href="https://useblume.dev/docs/cli/audit">
no X-Robots-Tag header
1A robots meta tag with noindex instead of a canonical comes from seo.noindex in frontmatter, or noindex: true on an archived version or an API reference source. A canonical pointing elsewhere comes from seo.canonical, from an archived version, whose pages point at their latest equivalent by default, or from an untranslated page that a locale fallback fills, which points at the page it copies.
For Google, inspect the page in Search Console's URL Inspection tool and select Test live URL. It fetches with Googlebot, which no spoofed request can stand in for, and reports Crawl allowed?, Page fetch, and Indexing allowed?. The indexed result shows the Google-selected canonical. If the page isn't indexed, see pages that are not indexed.
OpenAI has no inspection tool. It says sites blocking OAI-SearchBot can still appear as navigational links, and its publisher FAQ says keeping a page out takes noindex, which its crawler can only read if robots.txt lets it in.
Run blume audit against the deployment
blume audit checks every page at once. Build the deployed commit, then point it at your domain:
npx blume build
npx blume audit --url https://docs.acme.example --only indexability,robots,sitemap,networkFrom dist/ it reports noindex pages, canonicals that point at missing or redirecting pages, and robots.txt rules that block a sitemap page (BLUME_AUDIT_ROBOTS_DISALLOWS_INDEXABLE). With --url it fetches every page, robots.txt, and the sitemap from the live site, and reports error responses (BLUME_AUDIT_HTTP_4XX, BLUME_AUDIT_HTTP_5XX), an unreachable robots.txt or sitemap, and BLUME_AUDIT_ROBOTS_HEADER_CONFLICT for an X-Robots-Tag: noindex header on a page whose HTML is indexable.
It has two limits. It fetches as a generic HTTP client, not a crawler, so a clean run means your host serves the pages, not that crawlers get through. And its robots.txt check reads only the User-agent: * group, which is why check-robots.mjs exists. If your host supplies the site URL at deploy time, expect one BLUME_AUDIT_SITE_INFERRED_AT_DEPLOY.
Record what you observed
Access changes whenever someone edits a CDN rule, so keep dated evidence, and add each URL Inspection result by hand:
{
date -u +"%Y-%m-%dT%H:%M:%SZ"
node check-robots.mjs https://docs.acme.example / /quickstart
./check-access.sh https://docs.acme.example/quickstart "Install the Acme CLI"
} | tee -a crawler-access.logYour checks show what a crawler would get; your traffic shows what real crawlers got. Cloudflare's AI Crawl Control lists requests per AI crawler, and Vercel's AI Bots Ruleset in Log mode records AI bot traffic without blocking it. With access logs in nginx's default combined format, count a crawler's requests by IP and status:
grep 'OAI-SearchBot' access.log | awk '{ print $1, $9 }' | sort | uniq -c | sort -rnAnyone can send that user agent, so check each IP against OpenAI's published ranges:
// Check whether an IP from your logs is in OAI-SearchBot's published ranges.
// Usage: node check-ip.mjs 203.0.113.7
import { BlockList } from "node:net";
const ip = process.argv[2];
const family = ip?.includes(":") ? "ipv6" : "ipv4";
const response = await fetch("https://openai.com/searchbot.json");
const { prefixes } = await response.json();
const ranges = new BlockList();
for (const prefix of prefixes) {
const cidr = prefix.ipv4Prefix ?? prefix.ipv6Prefix;
const [address, bits] = cidr.split("/");
ranges.addSubnet(address, Number(bits), prefix.ipv4Prefix ? "ipv4" : "ipv6");
}
console.log(
ranges.check(ip, family)
? `${ip} is in OAI-SearchBot's published ranges`
: `${ip} is not in OAI-SearchBot's published ranges: treat it as unverified`
);For Googlebot, Google's check is a reverse DNS lookup ending in googlebot.com, google.com, or googleusercontent.com, then a forward lookup returning the same IP:
$ host 66.249.66.1
1.66.249.66.in-addr.arpa domain name pointer crawl-66-249-66-1.googlebot.com.
$ host crawl-66-249-66-1.googlebot.com
crawl-66-249-66-1.googlebot.com has address 66.249.66.1Verified crawler requests that got a 200 are the strongest evidence of access you can collect. To see whether access turns into visits, read measuring AI search referrals.
Troubleshooting
The live robots.txt isn't Blume's generated file
If dist/robots.txt already differs, a public/robots.txt replaced it; Blume never overwrites one. If only the live file differs, check Cloudflare's managed robots.txt, which prepends groups disallowing training crawlers such as GPTBot, ClaudeBot, and Google-Extended. Rerun check-robots.mjs after any change there.
Your docs live under a subpath of another site
Crawlers only read robots.txt at a host's root, so on a subpath deploy Blume's /docs/robots.txt is ignored and the main site's file decides. Run check-robots.mjs against the root domain with your docs paths. The Next.js subpath guide covers that setup.
Crawlers get 403 or 503 while browsers get 200
That's a WAF rule or bot challenge, not Blume. If robots.txt itself is blocked, Google treats a 4xx other than 429 as no rules. A 5xx stops Google crawling the site for 12 hours; after that Google uses the last good copy for up to 30 days, then acts as if there were no robots.txt, or stops crawling a site that's still generally unavailable. OpenAI recommends allowing its published IP ranges as well as OAI-SearchBot in robots.txt.
Pages send X-Robots-Tag: noindex
The header overrides the HTML. Vercel adds it to preview deployments and to production deployments a newer one replaced. If your domain or a proxy points at one of those, point it at the current production deployment.
You fixed robots.txt, but ChatGPT search still leaves the site out
OpenAI says its search systems can take about 24 hours to adjust to a robots.txt change. After that, rerun check-robots.mjs and check-access.sh, and confirm the firewall isn't blocking OAI-SearchBot's IP ranges. Access is the part you control; which pages ChatGPT search shows is up to OpenAI.
Google indexes the page, but AI Overviews never link it
If URL Inspection shows the page indexed and nothing limits its snippet, it meets every technical requirement Google lists. Google also says meeting them doesn't guarantee a page is shown, so no access setting is left to change.
Next step
Audit your live site
Build the commit you deployed, then point the audit at your domain to catch error responses, noindex headers, and an unreachable robots.txt or sitemap.
npx blume audit --url https://docs.acme.exampleA step here not working for you? Report a broken step.