Search
Add Pagefind search to a static documentation site
Switch your docs search to Pagefind, build the index with the site, and check headings, languages, long pages, and exclusions with a query script you rerun on every build.
By Hayden Bleasel12 min read

You don't need a search server to search static docs. Blume's default engine, Orama, writes one index file during the build and searches it in the reader's browser. Pagefind is the other keyless option: with search: pagefind() in your config, blume build renders your pages, Pagefind indexes the finished HTML into small pieces in dist/pagefind/, and the browser downloads only the pieces each query needs.
By the end of this guide you have a small Acme docs site with representative content, built with Pagefind, and a query checklist you rerun after every build to confirm that headings, languages, long pages, and exclusions behave.
Pagefind isn't an upgrade for every site. If your docs are small, or you want search while you write in blume dev, keep the default. If you're also weighing hosted or semantic search, Choose local, hosted or semantic search compares them all.
When Pagefind fits better than the default
Both engines run in the browser, need no API keys, and ship as files in dist/. They differ in what they index and what a reader downloads:
| Orama (default) | Pagefind | |
|---|---|---|
| Indexes | Your source files | The built HTML, after the build |
| What a search downloads | The whole blume-search.json, on first open | The index chunks for the words typed, plus one file per result shown |
In blume dev | Yes, live as you edit | No |
| Code blocks | Only with indexing.includeCodeBlocks | Always |
| Languages | One tokenizer, chosen by the default locale | A separate index for each page language |
| Versioned docs | Scoped to the version being read, with an "All versions" toggle | Every version at once, with no toggle |
| Section filters and breadcrumbs | Yes | No |
That makes Pagefind the better fit in three cases:
- The index file gets heavy. Orama's single file grows with every page, and each reader downloads all of it the first time they open search. Pagefind fetches only the chunks a query needs.
- Translations use a different script from your default language. Orama picks one tokenizer from
i18n.defaultLocale, so on an English-default site, Japanese, Chinese, Korean, or Russian translations aren't searchable. Pagefind indexes each language on its own. - Readers search for code. Option names and error codes often live only in code blocks, which Pagefind always indexes.
Before you switch an existing site, measure what you have. Build it once with the default and check the size of the one file every searching reader downloads:
npx blume build
ls -lh dist/blume-search.jsonCreate a project with representative content
Search problems hide in content a hello-world page doesn't have: headings, tabs, code, long pages, pages you don't want found, and other languages. Build a small site with all of them. Scaffold a project:
npx blume init acme-docs --yes
cd acme-docsReplace the starter page with a short introduction:
---
title: Acme docs
description: Send transactional email and SMS with the Acme Messages API.
---
Acme sends transactional email and SMS through one API. Start with
[Send a message](/guides/send-a-message), then look up any command in the
[CLI reference](/reference/cli).Add a guide with a heading, content inside tabs, and a term that only appears in a code block:
---
title: Send a message
description: Send an email or SMS with one request, and retry it safely.
---
Every message goes through `POST /messages`. Pick a channel first.
<Tabs>
<Tab title="Email">
Set `channel` to `email` and pass a `subject`. Acme renders your template
and sends it from your verified domain.
</Tab>
<Tab title="SMS">
Set `channel` to `sms`. A message longer than 160 characters is split into
segments, and each segment counts against your quota.
</Tab>
</Tabs>
## Send the request
```bash
curl https://api.acme.example/v1/messages \
-H "Authorization: Bearer $ACME_API_KEY" \
-H "Idempotency-Key: order-1234-receipt" \
-d channel=email -d to=ada@example.com -d template=receipt
```
## Retries
If a request times out, send it again with the same idempotency key. Acme
returns the first message instead of sending a second one.Add two pages that should never appear in search: one hidden from the sidebar, and one excluded outright.
---
title: On-call runbook
description: What to check when message delivery slows down.
sidebar:
hidden: true
---
Page the on-call engineer when the delivery backlog passes 10,000 messages,
then drain the dead-letter queue.---
title: Webhooks (v1)
description: The retired v1 webhook format, kept for existing integrations.
search:
exclude: true
---
v1 webhooks carry a signature header computed with your v1 signing secret.
Move to the current format before the v1 endpoint shuts down.Translate the guide into Japanese, and leave every other page untranslated:
---
title: メッセージを送信する
description: 1 回のリクエストでメールまたは SMS を送信します。
---
すべてのメッセージは `POST /messages` で送信します。リクエストがタイムアウトした場合は、同じ冪等キーを付けて再送信してください。Last, you need a long page, like the CLI or configuration reference most real sites have. This script writes one with a section for each of 64 commands, and puts an error code that appears nowhere else in its last section:
// Writes docs/reference/cli.mdx: one long page with a section per command.
import { mkdirSync, writeFileSync } from "node:fs";
const resources = ["messages", "templates", "domains", "webhooks", "keys", "logs", "suppressions", "workspaces"];
const verbs = ["list", "get", "create", "update", "delete", "export", "import", "watch"];
const lines = [
"---",
"title: CLI reference",
"description: Every acme command, its flags, and what it prints.",
"---",
"",
];
for (const resource of resources) {
for (const verb of verbs) {
lines.push(
"## acme " + resource + " " + verb,
"",
"Runs `" + verb + "` against " + resource + " in the current workspace. " +
"Pass `--workspace` to target another workspace, and `--json` to print " +
"the API response instead of a table. The command reads your key from " +
"`ACME_API_KEY` and exits with a non-zero code when the request fails.",
"",
"| Flag | Default | Description |",
"| --- | --- | --- |",
"| `--workspace` | current | The workspace to act on. |",
"| `--json` | `false` | Print the API response as JSON. |",
"| `--limit` | `50` | The most rows to print. |",
""
);
}
}
// One term that appears nowhere else on the site, deep in the page.
lines.push(
"## Exit codes",
"",
"Exports count against your daily quota. Past it, every command exits with " +
"`ACME_QUOTA_EXCEEDED` until the quota resets at midnight UTC.",
""
);
mkdirSync("docs/reference", { recursive: true });
writeFileSync("docs/reference/cli.mdx", lines.join("\n"));node scripts/cli-reference.mjsThe page comes out at about 6,000 words. It and the guide each mention "quota" once, which shows how Pagefind ranks a long page against a short one.
Switch search to Pagefind
Replace the config with one that adds the Japanese locale and switches the search adapter:
import { defineConfig } from "blume";
import { pagefind } from "blume/search";
export default defineConfig({
title: "Acme",
description: "Documentation for the Acme Messages API.",
i18n: {
defaultLocale: "en",
locales: [
{ code: "en", label: "English" },
{ code: "ja", label: "日本語" },
],
},
search: pagefind(),
});pagefind() takes no options, and Pagefind ships with Blume, so there's nothing else to install. To set popular pages or indexing options too, use the object form and pass the adapter as provider: pagefind().
Build and preview the site
Pagefind runs at the end of blume build, so there's no index in blume dev. Open search there and the dialog says "Search is available in the production build." Build instead:
npx blume buildAfter the pages render, the log prints Building search index and then Indexed N page(s) for search. That count covers every HTML file Pagefind read, including the ones it skipped, so it runs higher than what's searchable. The real counts are per language:
cat dist/pagefind/pagefind-entry.jsonUnder languages, en should have a page_count of 3 (the introduction, the guide, and the CLI reference) and ja a count of 1. Beside it are pagefind.js, per-language metadata and Wasm files, and the index/ and fragment/ chunks. Pagefind's own UI files, like pagefind-ui.js, are there too; Blume's dialog doesn't use them.
Now serve the build:
npx blume previewOpen http://localhost:4321, press ⌘K (or Ctrl K), and type retries. You get up to 12 results, each with the page title and Pagefind's excerpt, matches highlighted. Selecting one opens a Not Found page in the preview, because Pagefind's links end in a slash (see Troubleshooting). Keep the preview running.
Run the query checklist
To repeat checks after every content change, script them. Each check is a query and the page that should rank first, or no page at all:
| Query | Expected first result | What it checks |
|---|---|---|
retries | Send a message | Heading text is indexed |
idempotency | Send a message | Code blocks are indexed |
segments | Send a message | Content in a tab that isn't open is indexed |
ACME_QUOTA_EXCEEDED | CLI reference | A term deep in a long page is found |
quota | Send a message | A short page outranks a long one for the same single mention |
backlog | None | The hidden page is left out |
signing | None | The excluded page is left out |
送信 (on English pages) | None | Japanese pages stay in their own index |
送信 (on Japanese pages) | メッセージを送信する | The Japanese index works |
quota (on Japanese pages) | None | Untranslated fallback pages are left out |
The script runs those queries through the same pagefind.js the dialog loads, read from dist/, against the index the preview server serves. pagefind.js expects a browser page and picks its language index from the page's <html lang>, so the script stands in a minimal document for it:
// Runs each query through the Pagefind bundle in dist/, against the site that
// `blume preview` serves, and fails when a page ranks in the wrong place.
import { resolve } from "node:path";
import { pathToFileURL } from "node:url";
const site = process.env.SITE ?? "http://localhost:4321";
const lang = process.argv[2] ?? "en";
// pagefind.js picks its language index from the page's <html lang>.
globalThis.document = {
currentScript: null,
querySelector: () => ({ getAttribute: () => lang }),
};
// `first` is the page that should rank first; null means no results at all.
const checks = {
en: [
{ query: "retries", first: "/guides/send-a-message/" },
{ query: "idempotency", first: "/guides/send-a-message/" },
{ query: "segments", first: "/guides/send-a-message/" },
{ query: "ACME_QUOTA_EXCEEDED", first: "/reference/cli/" },
{ query: "quota", first: "/guides/send-a-message/" },
{ query: "backlog", first: null },
{ query: "signing", first: null },
{ query: "送信", first: null },
],
ja: [
{ query: "送信", first: "/ja/guides/send-a-message/" },
{ query: "quota", first: null },
],
};
const pagefind = await import(
pathToFileURL(resolve("dist/pagefind/pagefind.js")).href
);
await pagefind.options({ basePath: `${site}/pagefind/`, baseUrl: "/" });
await pagefind.init();
let failures = 0;
for (const { query, first } of checks[lang]) {
const { results } = await pagefind.search(query);
const top = results[0] ? (await results[0].data()).url : null;
const pass = top === first;
if (!pass) failures += 1;
console.log(
`${pass ? "pass" : "FAIL"} ${query} -> ${top ?? "no results"} (${results.length} results)`
);
}
await pagefind.destroy();
process.exitCode = failures > 0 ? 1 : 0;Run it once per language from the project root, in a second terminal while blume preview is still serving:
node search-check.mjs en
node search-check.mjs jaEvery line should start with pass, and the script exits with an error when any check fails. Pagefind's result URLs end in a slash, because each page is a folder with an index.html, so write your expected URLs the same way.
Headings
retries finds the guide through its ## Retries heading, and Pagefind weights heading text above body text. Pagefind also tracks which heading each match sits under, but Blume's dialog lists one row per page. Selecting a result opens the page at its top, and the excerpt shows the passage that matched.
Language selection
Pagefind builds one index per <html lang>, and the browser loads the index for the language of the page the reader is on. On a /ja/ page, search covers Japanese pages only; on an English page, English only. The dialog still shows its "All languages" toggle, but it doesn't widen a Pagefind search. Blume loads Pagefind once per full page load, and Pagefind reads the page's language as it loads, so after you switch language from the language menu, reload the page before you test search. Pagefind's npm package ships its extended build, which splits Chinese, Japanese, and Korean text into words, and Make docs search work in Chinese, Japanese and Korean goes deeper on those languages.
Long pages
Pagefind compares each page's length with the average on your site, so a term mentioned once on a 6,000-word page counts for less than the same term once on a 100-word page. That's the quota check. The long page is still found for terms only it contains, and its excerpt comes from the passage with the most matches, but the reader lands at the top and has to find the passage on the page.
If a long page should win for its own terms, give it search.boost or search.keywords in its frontmatter. Blume turns a boost into a Pagefind weight, capped at 10, and Pagefind damps repeated matches, so the lift is smaller than the number suggests. Rerun the checklist after either change. A page that covers many unrelated topics is often better split.
Inspect what's excluded
Blume marks up every docs page for Pagefind. The page's article carries data-pagefind-body, so Pagefind indexes only the article, never the header, sidebar, table of contents, or footer. Once any page has that attribute, Pagefind skips every page that doesn't. A page that shouldn't be searched gets data-pagefind-ignore="all" on its <html> element instead. List both kinds in your build:
# Every HTML page Pagefind leaves out
grep -rL "data-pagefind-body" dist --include='*.html'
# The ones Blume excluded on purpose
grep -rl 'data-pagefind-ignore="all"' dist --include='*.html'For the Acme site, the first list includes 404.html, the runbook, the v1 webhooks page, and every untranslated page under /ja/. What's left out, and what brings it back:
| Left out | Why | To include it |
|---|---|---|
A page with search.exclude: true | Its frontmatter says so | Remove the setting |
A page with sidebar.hidden: true | Hidden pages are excluded by default | Set indexing: { includeHiddenPages: true } in the search object |
| An untranslated page under a locale | It repeats the default language's content | Translate the page |
Custom pages built on PageLayout, the 404 page, and the generated changelog index | Blume doesn't mark them as docs content | Move content you want found into a docs page |
An API reference embedded with scalar() | It draws in the browser, so the built page has no text to index | Use openapi() or asyncapi(), which build a page per operation |
Inside an article, Pagefind indexes everything in the HTML, including tab panels that aren't open, except the elements it skips by default, like <nav>, <footer>, <script>, and <form>. The MCP server and the assistant keep their own indexes, so Pagefind doesn't change what they find.
Deploy and check production
dist/pagefind/ deploys with the rest of dist/ to any static host, with nothing else to run. Once it's live, run the checklist against your domain. The script still loads pagefind.js from your local dist/, so build the same commit first, and include any base path in SITE:
SITE=https://docs.acme.example node search-check.mjs en
SITE=https://docs.acme.example node search-check.mjs jaThen select a result on the live site, to confirm your host serves the page at the URL Pagefind links to.
Troubleshooting
Search says it's available in the production build
You're in blume dev, which has no Pagefind index. Test search with blume build and blume preview, or keep Orama if you need search while writing.
Search says something went wrong
"Something went wrong. Please try again." in a production build means the dialog couldn't load Pagefind's files. Check that the build log has the Building search index line, that dist/pagefind/pagefind.js exists, and that the same path loads on your site, under any base path. After an upgrade that changes Pagefind's version, a browser can pair a cached pagefind.js with the new index, and the console warns "Pagefind JS version doesn't match the version in your search index." Reloading without the cache fixes it.
A result opens a Not Found page in blume preview
Pagefind links to each page's folder URL, with a trailing slash, like /guides/send-a-message/. Blume sites use slashless URLs, and the preview server answers a slashed one with "Not Found (trailingSlash is set to "never")". Remove the slash from the address bar to confirm the page exists, and check result links on your deployed site, where it depends on whether your host serves the folder's index.html at the slashed URL or redirects it to the slashless one.
Results come from an old docs version
Pagefind doesn't scope results by version, so on a versioned site a search covers every snapshot, and the dialog hides the "All versions" toggle. If readers need results from the version they're reading, stay on Orama, which scopes them.
An exclusion check fails on a page without the word
Pagefind matches the start of words, so retr finds "Retries", and it can match part of a longer word you type: in testing, runbook matched a page containing "Runs". For checks that expect no results, pick a word that doesn't start with, or contain, any other word on your site.
Next step
Switch your site to Pagefind
Set search to pagefind() in your config, build, and run the query checklist against blume preview.
npx blume buildA step here not working for you? Report a broken step.