Skip to content
Blume
Esc
↑↓navigate↵open⌘Jpreview
Guides

Discoverability

Find out why Google isn't indexing your docs pages

Read the reason Search Console gives, trace robots.txt, noindex, and canonical causes to the file behind them, and confirm each fix on the live site.

By 8 min read

When Search Console lists a docs page as not indexed, the reason usually comes down to one of four causes. Google can't crawl the page (a robots.txt rule, an error status, or a login). The page tells Google not to index it (a noindex tag or an X-Robots-Tag header). Google indexed a different URL as the canonical. Or Google found or crawled the page and chose not to index it for now. URL Inspection tells you which one applies, and blume audit traces the first three to the file or setting behind them.

This guide runs that diagnosis on an example Acme docs site at docs.acme.example, with one page broken on purpose for each cause and the audit finding it produces, then shows how to confirm a fix. The Search Console steps work for any site; the audit needs a Blume build.

Not every exclusion is a bug. Google says not to expect every URL to be indexed, and a noindex page, an archived version, a redirect, or a duplicate pointing at its original is out on purpose. Blume's defaults avoid the common accidents, so the cause is usually a setting someone added. This guide is about reading the report, not getting more pages crawled, and nothing here can make Google index a page or say when.

Inspect the exact URL

In Search Console, open URL Inspection and paste the complete URL as the report shows it. /quickstart, /quickstart/, and the http:// version are different URLs to Google, and Blume's canonical tags name the one without the trailing slash. Then expand Page indexing:

  • Discovery lists the sitemaps and referring page Google found the URL through.
  • Crawl: Crawl allowed? is about robots.txt, Page fetch about the HTTP response, and Indexing allowed? about noindex.
  • Indexing shows the User-declared canonical and the Google-selected canonical. If Google's pick isn't the URL you inspected, Google indexed that one instead.

This all comes from Google's last crawl. Test live URL fetches the page as deployed now, though it doesn't check duplicate or canonical conditions. When robots.txt blocks a page, Indexing allowed? always reads "Yes", because Google never saw a noindex.

Don't decide from a site: search. Google says it doesn't necessarily return every indexed URL, so a missing page isn't proof that it's out of the index.

Match the reason to a cause

The Page indexing report lists excluded URLs under "Why pages aren't indexed". Its Source column says whether a reason comes from your website or from Google, and Google says you can generally fix only the Website ones. The common reasons map to Blume like this (check ids without their BLUME_AUDIT_ prefix):

Search Console reasonUsual cause on a Blume siteAudit check
URL blocked by robots.txtA Disallow rule in your own public/robots.txtROBOTS_DISALLOWS_INDEXABLE
URL marked ‘noindex’noindex in frontmatter, a version, or a reference source, or an X-Robots-Tag header from the hostROBOTS_META_UNEXPECTED, ROBOTS_HEADER_CONFLICT
Alternate page with proper canonical tag; Duplicate, Google chose different canonical than user; Duplicate without user-selected canonicalseo.canonical, an archived version or translation fallback, near-duplicate pages, or a wrong deployment.siteCANONICAL_NOT_SELF, DUPLICATE_CONTENT
Not found (404), Server error (5xx), Blocked due to unauthorized request (401)A deleted or moved page, a host rewrite, or access protectionHTTP_4XX, HTTP_5XX
Page with redirectAn entry in redirects, as intendedNone
Discovered - currently not indexed; Crawled - currently not indexedGoogle's decisionORPHAN_PAGE, INDEXABLE_PAGE_NOT_IN_SITEMAP

Run the audit on the build

blume audit crawls the HTML in dist/ and names the source file (and frontmatter line) behind each finding. It needs an absolute site URL: without one, Blume writes no canonical tags or sitemap, and the audit reports BLUME_AUDIT_SITE_NOT_SET instead.

import { defineConfig } from "blume";

export default defineConfig({
  title: "Acme Docs",
  deployment: { site: "https://docs.acme.example" },
});

Build the same commit that's in production, then run the checks that decide indexing:

npx blume build
npx blume audit --only indexability,sitemap,robots,links,duplicates

Headers and status codes live on the deployed site, so check it too:

npx blume audit --url https://docs.acme.example --only indexability,network

Each section below shows what the audit reported for Acme's broken page.

Blocked crawling

Blume's own robots.txt allows every crawler, but Blume never overwrites a file you put in public/, so one left over from staging keeps shipping. Acme's file still blocks a page that was being rewritten:

User-agent: *
Disallow: /templates

Sitemap: https://docs.acme.example/sitemap.xml
✖ robots.txt disallows a page that is in the sitemap  1 page
    /templates                         dist/robots.txt
    fix: A page can't be both disallowed in robots.txt and advertised in the sitemap.

Remove the rule, or delete the file to get Blume's back. The audit only reads the rules for User-agent: *, so check any Googlebot group, and CDN bot rules, with the scripts in checking crawler access.

Don't use robots.txt to keep a page out of Google. Google can still index a blocked URL if other pages link to it ("Indexed, though blocked by robots.txt"), and the block hides any noindex on the page. Use noindex instead.

A noindex you didn't mean

Acme's webhooks page was noindex while the feature was in beta, and the line outlived the beta:

---
title: Webhooks
description: Receive delivery, bounce, and reply events from Acme at an HTTPS endpoint you control.
noindex: true
---
ℹ Page is not indexable  1 page
    /webhooks                          docs/webhooks.mdx:4
    fix: Remove `noindex` from the page's frontmatter if it should be indexed.

It's only a note, since noindex is usually deliberate, so read every page in this group. The same tag comes from seo.noindex, from noindex: true on an archived version, and from noindex: true on an API reference source.

A page's HTML can be clean while the host sends X-Robots-Tag: noindex, and the header wins. Only the live check sees it:

✖ X-Robots-Tag header conflicts with the page's robots meta  5 pages
    /getting-started                   docs/getting-started.mdx
    /                                  docs/index.mdx
    /quickstart                        docs/quickstart.mdx
    … and 2 more (--verbose)
    fix: Remove the X-Robots-Tag header, or align it with the page's robots meta.

Look for a header rule in your host or CDN config. On a preview deployment it's expected: Vercel adds X-Robots-Tag: noindex to every preview, and to production deployments a newer one has replaced.

Canonical consolidation

Google indexes one URL per set of duplicates. Blume gives every indexable page a canonical tag naming its own URL, built from deployment.site. It names another page in three cases: a page's own seo.canonical, an archived version's page pointing at the live one, and a translation fallback pointing at the page it copies. The last two show up as "Alternate page with proper canonical tag", which Google says needs no action.

Acme kept an old page and pointed it at the quickstart:

---
title: Getting started
description: The steps from the quickstart, kept at their old URL for readers who bookmarked it.
seo:
  canonical: https://docs.acme.example/quickstart
---
✖ Non-canonical page in sitemap  1 page
    /getting-started                   docs/getting-started.mdx:5
    fix: List only canonical URLs in the sitemap.

ℹ Non-canonical page  1 page
    /getting-started                   docs/getting-started.mdx:5
    fix: Point `seo.canonical` at this page, or remove it to use the default self-canonical.

Blume's generated sitemap still lists a page whose seo.canonical points elsewhere, and the audit reports that as an error. For a page that really is a duplicate, a redirect is cleaner. Delete the file and add this to blume.config.ts:

redirects: [{ from: "/getting-started", to: "/quickstart", status: 301 }],

The old URL then moves to "Page with redirect", as it should. For "Duplicate, Google chose different canonical than user", compare the two canonicals in URL Inspection. If Google's is on another host, check the tag on the live page:

curl -s https://docs.acme.example/quickstart | grep -o '<link[^>]*rel="canonical"[^>]*>'

Blume builds its canonicals from deployment.site. For "Duplicate without user-selected canonical", note that DUPLICATE_CONTENT only catches byte-identical pages. Its fix text suggests seo.canonical, which runs into the sitemap error above, so merge the pages and redirect the one you drop.

Discovered or crawled, but not indexed

These are Google's calls. "Crawled - currently not indexed" means Google fetched the page and left it out; it may or may not be indexed later, and Google says there's no need to resubmit it. "Discovered - currently not indexed" means Google hasn't crawled the URL yet, typically because it expected the crawl to overload the site. No setting clears either, but three things are worth a look.

The content is in the HTML. Blume prerenders every docs page, so confirm the deployed page answers 200 and carries a phrase from its text, and compare with View crawled page in URL Inspection:

curl -sI https://docs.acme.example/quickstart | head -n 1
curl -s https://docs.acme.example/quickstart | grep -c "API key"

Something links to it. Google asks that every page you care about have a link from at least one other page on your site. Acme's beta page is hidden from the sidebar:

---
title: Scheduled sends
description: Schedule a message to send later with the sendAt field, available in beta on selected accounts.
hidden: true
---
⚠ Orphan page (only reachable from navigation)  1 page
    /scheduled-sends                   docs/scheduled-sends.mdx
    fix: Link to this page from the body of a related page.

⚠ Indexable page not in sitemap  1 page
    /scheduled-sends                   docs/scheduled-sends.mdx
    fix: Remove `draft`/`hidden`/`noindex` from the page's frontmatter if it should be indexed.

hidden drops a page from the sidebar, pagination, and sitemap, but doesn't add noindex. The page is live and indexable with nothing pointing at it, and URL Inspection may call it unknown to Google. Link it from a related page if readers should find it, or add noindex: true if they shouldn't.

The page stands on its own. For crawled pages, look at DUPLICATE_TITLE, DUPLICATE_DESCRIPTION, and LOW_WORD_COUNT. For discovered ones, the live check reports slow responses, timeouts, and 5xx errors. Don't answer either reason by publishing more pages or resubmitting the same ones.

Verify the fix and follow the recrawl

  1. Rebuild and rerun the audit until the finding is gone. Deploy, then rerun it with --url.
  2. In URL Inspection, run Test live URL. Crawl allowed? and Indexing allowed? should both read "Yes".
  3. For a page that matters, click Request Indexing. Google refuses it while the live test finds the page non-indexable.
  4. Once every instance of a reason is fixed, open it in the report and click Validate fix. Google emails you when validation passes or fails; don't click it again before then.

Google recrawls on its own schedule, and the report updates as it does, with or without validation.

Troubleshooting

The audit says deployment.site is inferred at deploy time

With the vercel() or netlify() adapter, Blume reads the site URL from the platform while it deploys. A local build has none, so the canonical and sitemap checks stay quiet and the audit reports BLUME_AUDIT_SITE_INFERRED_AT_DEPLOY. Don't hardcode the site. Build with the platform's variables set instead, as the finding suggests (on Vercel, VERCEL=1 VERCEL_PROJECT_PRODUCTION_URL=docs.acme.example npx blume build), and use --url for the live checks.

Google chose a canonical on another domain

deployment.site points at the wrong origin, like an old domain or a per-deploy URL. Cloudflare Pages only exposes each deployment's own URL, so set site there yourself (see Set your site URL).

The audit is clean, but the report still lists the page

The report reflects the last crawl. If URL Inspection's Last crawl date is older than your fix, the report catches up when Google recrawls.

Old URLs show as Not found (404)

Google keeps trying URLs it once knew. A 404 for a page you removed isn't necessarily a problem, but a page that moved needs a redirect. See Move documentation URLs while preserving old links.

Next step

Audit your build

After your next build, run the checks this guide leans on. Add --url with your domain to include the live site.

npx blume audit --only indexability,sitemap,robots,links
Read the audit reference

A step here not working for you? Report a broken step.

Keep going.More guides.

Upgrade your docs with Blume.

Install today and ship a production-grade docs site in minutes. Free and open source, forever.

npx blume init