Search
Choose local, hosted or semantic search for your docs
Run one fixed query set through local, hosted, and semantic search, compare what each finds and what it takes to run, and keep the engine your docs need.
By Hayden Bleasel10 min read

For most documentation, the best search is the one Blume already ships: a local Orama index with no keys or service, working in blume dev. Switch only when a fixed set of real queries shows it missing pages another engine finds. Pagefind is the other local option, built from your rendered HTML. Hosted keyword search (Algolia, Typesense, Orama Cloud) moves the index to a service Blume uploads to on every build. Semantic search (Mixedbread) matches meaning instead of words, but needs server output and an upload step you run yourself.
You end with a query set, a script that runs it through the real search dialog of each candidate build, a results file, and a choice that follows from it. We ran the same method on Blume's own docs, and the numbers are below. It covers the search dialog only, since the assistant and the MCP server search their own index. It doesn't compare prices.
Local, hosted, and semantic search in Blume
Every adapter comes from blume/search and plugs into the same dialog, so switching is one line of config. What changes is where the index lives and what it can filter:
| Adapter | Index | In blume dev | Language | Version | Also needs |
|---|---|---|---|---|---|
orama(), the default | Local, built from your source | Yes | Scoped | Scoped | Nothing |
flexsearch() | Local, built from your source | Yes | Scoped | Scoped | The flexsearch package |
pagefind() | Local, built from the HTML, loaded in pieces | No | The page's language | All versions | Nothing |
algolia() | Hosted, uploaded on each build | Last uploaded index | Scoped | Scoped | algoliasearch, ALGOLIA_ADMIN_API_KEY |
typesense() | Hosted, uploaded on each build | Last uploaded index | Scoped | Scoped | typesense, TYPESENSE_ADMIN_API_KEY |
oramaCloud() | Hosted, uploaded on each build | Last uploaded index | Scoped | All versions | @oramacloud/client, ORAMA_PRIVATE_API_KEY, an indexId |
mixedbread() | Semantic, filled by your own sync | Through /api/search | All languages | All versions | @mixedbread/sdk, MIXEDBREAD_API_KEY, server output |
What gets indexed differs too. Indexes built from your source strip code blocks unless you set search.indexing.includeCodeBlocks, Pagefind indexes each page's rendered article with its code, and Mixedbread searches the raw Markdown you upload. The search.keywords and search.boost frontmatter fields reach every adapter except Mixedbread (see What's indexed).
Write a fixed query set
A comparison is only as good as its queries, so take them from what readers type. With an analytics adapter configured, every search is a search event, and the ones with results: 0 are your gaps (see Find missing documentation from searches with no results). Add questions from support tickets. Cover each kind of query, because engines fail on different ones:
- Exact terms, like a config key or an error name.
- Terms that appear only inside code blocks.
- Questions typed as sentences.
- Other words for what a page covers.
- Typos.
- A query from a translated page and from an older version, if your docs have them.
Aim for 20 to 40 queries, list every page that correctly answers each, then freeze the file. These are the 12 we used on Blume's own docs:
[
{ "query": "rate limiting", "kind": "exact term", "expect": ["/docs/configuration/rate-limiting"] },
{ "query": "llms.txt", "kind": "exact term", "expect": ["/docs/discoverability/llms-txt"] },
{ "query": "includeCodeBlocks", "kind": "only in code", "expect": ["/docs/configuration/search"] },
{ "query": "how do I hide a page from search", "kind": "question", "expect": ["/docs/configuration/search", "/docs/content/navigation"] },
{ "query": "can I host the docs on my own server", "kind": "question", "expect": ["/docs/deployment"] },
{ "query": "how do I translate my docs", "kind": "question", "expect": ["/docs/content/i18n", "/docs/cli/translate"] },
{ "query": "chatbot", "kind": "other words", "expect": ["/docs/configuration/assistant"] },
{ "query": "dark mode", "kind": "other words", "expect": ["/docs/configuration/theming"] },
{ "query": "keep old links working", "kind": "other words", "expect": ["/docs/deployment"] },
{ "query": "deploymnet", "kind": "typo", "expect": ["/docs/deployment"] },
{ "query": "frontmater", "kind": "typo", "expect": ["/docs/content/frontmatter"] },
{ "query": "sidebar order", "kind": "exact term", "expect": ["/docs/content/meta", "/docs/content/navigation"] }
]Run the set through each candidate
Test the dialog rather than each provider's API, since Blume adds its own filters and ranking. This script opens a built site in headless Chromium, types each query into the search dialog, and appends where the first correct page ranked to results.csv. It reads the dialog's internal markup, which isn't a public API, so check it still works after you upgrade Blume.
npm install --save-dev playwright@1.63.0
npx playwright install chromium// Runs every query in search-queries.json through a Blume site's search
// dialog and records where the first correct page ranks.
// Usage: node search-eval.mjs <site-url> <candidate-name>
import { existsSync } from "node:fs";
import { appendFile, readFile } from "node:fs/promises";
import { chromium } from "playwright";
const [site, candidate] = process.argv.slice(2);
const queries = JSON.parse(await readFile("search-queries.json", "utf8"));
const date = new Date().toLocaleDateString("en-CA");
const pathOf = (href) => {
const path = new URL(href).pathname;
return path.length > 1 && path.endsWith("/") ? path.slice(0, -1) : path;
};
if (!existsSync("results.csv")) {
await appendFile("results.csv", "date,candidate,query,kind,rank,results\n");
}
const browser = await chromium.launch();
const page = await browser.newPage();
await page.goto(site);
await page.keyboard.press("/");
const input = page.locator("[data-blume-search-input]");
await input.waitFor();
let top1 = 0;
let top3 = 0;
for (const { query, kind, expect } of queries) {
await input.fill(query);
// Settled once a row renders, or a message other than the loading "…".
await page.waitForFunction(() => {
const row = document.querySelector("#blume-search-listbox [role=option]");
const note = document.querySelector("[data-blume-search-message]");
return row || (note && !note.hidden && note.textContent !== "…");
});
const urls = await page.$$eval("#blume-search-listbox a[role=option]", (links) =>
links.map((link) => link.href)
);
const rank = urls.map(pathOf).findIndex((url) => expect.includes(url)) + 1;
if (rank === 1) top1 += 1;
if (rank > 0 && rank <= 3) top3 += 1;
const cells = [date, candidate, JSON.stringify(query), kind, rank || "miss", urls.length];
await appendFile("results.csv", `${cells.join(",")}\n`);
}
await browser.close();
console.log(`${candidate}: top 1 ${top1}/${queries.length}, top 3 ${top3}/${queries.length}`);Build each candidate, changing only search in blume.config.ts, and serve it with blume preview, which listens on port 4321. Start with the default, since it's your baseline:
npx blume build
npx blume preview
# in a second terminal:
node search-eval.mjs http://localhost:4321 oramaStop the preview before each new build. For Pagefind, set search: pagefind() and repeat. Always test a build, since Pagefind's index doesn't exist in blume dev. Pagefind's result links end in a slash, like /docs/deployment/, which blume preview answers with Not Found. The script strips the slash before it compares, so the ranks still count; the Pagefind guide covers what that means on your host.
A hosted keyword candidate
Pick one provider to stand for the category. With Algolia, create an index, put the search-only key in the config, and put the admin key in .env.local, which Blume loads for every command:
import { defineConfig } from "blume";
import { algolia } from "blume/search";
export default defineConfig({
search: algolia({
appId: "YOUR_APP_ID",
apiKey: "YOUR_SEARCH_ONLY_KEY",
indexName: "docs_eval",
}),
});npm install algoliasearch
echo "ALGOLIA_ADMIN_API_KEY=your-admin-key" >> .env.local
npx blume build
npx blume previewThe build log should say Synced 100 record(s) to algolia, with your page count. If it says Search sync skipped: instead, nothing was uploaded and the results don't count.
A semantic candidate
Mixedbread's queries go through a /api/search route on your server, so this candidate needs a host adapter. Use node(), since blume preview can serve its build locally:
import { defineConfig } from "blume";
import { node } from "blume/deploy";
import { mixedbread } from "blume/search";
export default defineConfig({
deployment: node(),
search: mixedbread({ storeId: "acme-docs" }),
});Blume doesn't upload anything to Mixedbread. You fill the store with Mixedbread's CLI, which reads MXBAI_API_KEY, while Blume's route reads MIXEDBREAD_API_KEY:
npm install @mixedbread/sdk
npm install --save-dev @mixedbread/cli@2.4.0
export MXBAI_API_KEY=your-mixedbread-key
npx mxbai store create acme-docs
npx mxbai store sync acme-docs "docs/**" --yes
echo "MIXEDBREAD_API_KEY=your-mixedbread-key" >> .env.local
npx blume build
npx blume previewCompare more than rank
The top 1 and top 3 counts are the headline, but read the misses by kind: an engine that only misses typos has a different problem from one that misses every question. Then check what the counts hide.
- Links. Blume's Mixedbread route takes each result's link from a
urlin the metadata Mixedbread generated for the file, which neithermxbai store syncnor Blume sets. A result without one links back to the current page, and the script counts it as a miss. Check what your store returns withnpx mxbai store search acme-docs "dark mode" --return-metadata --format jsonbefore you trust semantic results. - Scoping. Run the script again from a translated page and from a page in an older version, and check that results stay in that language and version where the table says they should. The "All languages" toggle shows on every adapter, but changes nothing on Pagefind, which keeps to the page's language, or Mixedbread, which searches every language.
- Freshness. A local index is rebuilt with the site. A hosted index holds whatever the last successful upload sent, and a failed upload only warns. Typesense is the exception: Blume deletes the collection before it uploads, so searches fail while a sync runs, and a sync that fails partway leaves the collection missing or partial. The Mixedbread store changes only when someone runs
mxbai store sync, so it needs its own CI step.
What we measured on Blume's docs
We ran the 12 queries on 2026-09-27 against Blume's own docs on the main branch: 100 English pages, docs plus changelog. Orama, FlexSearch, and Typesense ran through Blume's own browser search clients, with Typesense 30.2 running locally and filled by Blume's upload, and Pagefind 1.5.2 indexed the pages as the dev server rendered them. The script above gave the same Orama row through the dialog. Mixedbread wasn't measured: it needs an account.
| Engine | Top 1 | Top 3 | Missed |
|---|---|---|---|
| Orama, the default | 4 | 5 | 6 |
| FlexSearch | 5 | 5 | 6 |
| Pagefind | 8 | 8 | 3 |
| Typesense | 2 | 3 | 6 |
- Questions were the weak spot everywhere: no engine found the deployment page for
can I host the docs on my own server. - Only Pagefind found
includeCodeBlocks, which appears only inside a code block, and only Pagefind put the right page first for both typos. Orama found nothing for either. - Adding
chatbotto the assistant page'ssearch.keywordsmoved it to first on Orama. - Typesense's three fifth-place results come from Blume's sort: relevance in ten bands, then
search.boost. With every boost at 1, thellms.txtpage landed behind four weaker matches. Sorted by relevance alone, the same index put 6 of 12 first. - The Orama index ships whole to each reader on first search, every language in one file: 3.0 MB for 300 pages in five languages, about 870 KB with gzip. Pagefind's first query loaded about 210 KB, plus its 45 KB script.
- With English as the default locale, the Japanese and Hindi pages returned nothing on Orama, even for their own titles. German worked.
- As Algolia records, 23 of the 100 pages are over 10 KB. Algolia caps record size by plan, from 10 KB to 100 KB (checked 2026-09-27), and Blume uploads one record per page.
These are observations from one site and 12 queries, not a ranking of engines. Pagefind's lead here came mostly from a code term and two typos, and one keyword closed one of Orama's misses. Run your own set before you switch.
Which search fits your docs
- Most docs: keep Orama. Fix misses with
search.keywords,search.boost, andincludeCodeBlocksfirst. Keywords and boosts also rank pages for the assistant and the MCP server. - Large or many-language static sites, or terms that live in code: Pagefind. You give up search in
blume devand version scoping. See Add Pagefind search. - Versioned docs: Orama, FlexSearch, Algolia, or Typesense. The others search every version at once.
- Pages in a non-Latin script: with a non-Latin default language, Orama or Pagefind, not FlexSearch. With translations in another script than the default, like Japanese under English, Orama drops their words: use Pagefind, or Algolia with the index's languages set. The assistant can't search those words either way. See Make docs search work in Chinese, Japanese and Korean.
- An index other tools also query, or ranking rules you manage at the provider: hosted keyword. Algolia keeps its index settings across uploads. Blume deletes and recreates the Typesense collection on every build, so settings tuned there need reapplying and searches fail while the sync runs (see Self-host documentation search with Typesense). Make sure someone sees the build's upload warning.
- Readers ask in their own words and every keyword engine misses: semantic search, or the assistant. Both need server output (see Static or server-rendered documentation). With the assistant on, the dialog already offers every query to it as an Assistant row.
The assistant and MCP server keep their own index
Your search adapter changes the dialog, nothing else. The assistant grounds each answer in pages it retrieves from an Orama index baked into the build, whatever the adapter, even with search: false. The MCP server's search_docs tool and the JSON API's search use the same kind of index. So switching to Mixedbread doesn't make the assistant semantic, and a better dialog score says nothing about its answers. Only the inkeep() assistant adapter brings its own retrieval. To test answers instead of result lists, see Test whether your documentation can answer user questions.
Troubleshooting
Every query misses on a hosted candidate
Open the dialog yourself. "Something went wrong. Please try again." means the request to the provider failed: check the adapter's public key and host or endpoint. Mixedbread fails differently: when /api/search errors, the dialog shows no results instead of an error. Read the terminal running blume preview, and check that MIXEDBREAD_API_KEY is set; a build without it warns with BLUME_MISSING_SECRET.
The hosted index is missing new pages
Look for Search sync skipped: in the build log. It names the cause: an unset admin key, an oramaCloud() without indexId, or the provider rejecting the upload, such as an Algolia record over your plan's size limit. The build still succeeds, so the dialog keeps serving whatever the provider already holds. On Typesense, that can be a missing or partial collection, so build again once the cause is fixed.
The build fails with BLUME_SERVER_FEATURE_REQUIRED
mixedbread() on a static build fails with Search (mixedbread) requires server output. Set a host adapter from blume/deploy, like node().
The build fails with BLUME_DEPENDENCY_MISSING
The FlexSearch, Algolia, Typesense, Orama Cloud, and Mixedbread SDKs don't ship with Blume. Install the package the message names, like npm install algoliasearch.
Next step
Find your first queries
Pull the searches that found nothing from your analytics, then run them through your current search before you change anything.
Read the search analytics guideA step here not working for you? Report a broken step.