Skip to content
Blume
Esc
↑↓navigate↵open⌘Jpreview
Guides

Search

Choose local, hosted or semantic search for your docs

Run one fixed query set through local, hosted, and semantic search, compare what each finds and what it takes to run, and keep the engine your docs need.

By 10 min read

For most documentation, the best search is the one Blume already ships: a local Orama index with no keys or service, working in blume dev. Switch only when a fixed set of real queries shows it missing pages another engine finds. Pagefind is the other local option, built from your rendered HTML. Hosted keyword search (Algolia, Typesense, Orama Cloud) moves the index to a service Blume uploads to on every build. Semantic search (Mixedbread) matches meaning instead of words, but needs server output and an upload step you run yourself.

You end with a query set, a script that runs it through the real search dialog of each candidate build, a results file, and a choice that follows from it. We ran the same method on Blume's own docs, and the numbers are below. It covers the search dialog only, since the assistant and the MCP server search their own index. It doesn't compare prices.

Local, hosted, and semantic search in Blume

Every adapter comes from blume/search and plugs into the same dialog, so switching is one line of config. What changes is where the index lives and what it can filter:

AdapterIndexIn blume devLanguageVersionAlso needs
orama(), the defaultLocal, built from your sourceYesScopedScopedNothing
flexsearch()Local, built from your sourceYesScopedScopedThe flexsearch package
pagefind()Local, built from the HTML, loaded in piecesNoThe page's languageAll versionsNothing
algolia()Hosted, uploaded on each buildLast uploaded indexScopedScopedalgoliasearch, ALGOLIA_ADMIN_API_KEY
typesense()Hosted, uploaded on each buildLast uploaded indexScopedScopedtypesense, TYPESENSE_ADMIN_API_KEY
oramaCloud()Hosted, uploaded on each buildLast uploaded indexScopedAll versions@oramacloud/client, ORAMA_PRIVATE_API_KEY, an indexId
mixedbread()Semantic, filled by your own syncThrough /api/searchAll languagesAll versions@mixedbread/sdk, MIXEDBREAD_API_KEY, server output

What gets indexed differs too. Indexes built from your source strip code blocks unless you set search.indexing.includeCodeBlocks, Pagefind indexes each page's rendered article with its code, and Mixedbread searches the raw Markdown you upload. The search.keywords and search.boost frontmatter fields reach every adapter except Mixedbread (see What's indexed).

Write a fixed query set

A comparison is only as good as its queries, so take them from what readers type. With an analytics adapter configured, every search is a search event, and the ones with results: 0 are your gaps (see Find missing documentation from searches with no results). Add questions from support tickets. Cover each kind of query, because engines fail on different ones:

  • Exact terms, like a config key or an error name.
  • Terms that appear only inside code blocks.
  • Questions typed as sentences.
  • Other words for what a page covers.
  • Typos.
  • A query from a translated page and from an older version, if your docs have them.

Aim for 20 to 40 queries, list every page that correctly answers each, then freeze the file. These are the 12 we used on Blume's own docs:

[
  { "query": "rate limiting", "kind": "exact term", "expect": ["/docs/configuration/rate-limiting"] },
  { "query": "llms.txt", "kind": "exact term", "expect": ["/docs/discoverability/llms-txt"] },
  { "query": "includeCodeBlocks", "kind": "only in code", "expect": ["/docs/configuration/search"] },
  { "query": "how do I hide a page from search", "kind": "question", "expect": ["/docs/configuration/search", "/docs/content/navigation"] },
  { "query": "can I host the docs on my own server", "kind": "question", "expect": ["/docs/deployment"] },
  { "query": "how do I translate my docs", "kind": "question", "expect": ["/docs/content/i18n", "/docs/cli/translate"] },
  { "query": "chatbot", "kind": "other words", "expect": ["/docs/configuration/assistant"] },
  { "query": "dark mode", "kind": "other words", "expect": ["/docs/configuration/theming"] },
  { "query": "keep old links working", "kind": "other words", "expect": ["/docs/deployment"] },
  { "query": "deploymnet", "kind": "typo", "expect": ["/docs/deployment"] },
  { "query": "frontmater", "kind": "typo", "expect": ["/docs/content/frontmatter"] },
  { "query": "sidebar order", "kind": "exact term", "expect": ["/docs/content/meta", "/docs/content/navigation"] }
]

Run the set through each candidate

Test the dialog rather than each provider's API, since Blume adds its own filters and ranking. This script opens a built site in headless Chromium, types each query into the search dialog, and appends where the first correct page ranked to results.csv. It reads the dialog's internal markup, which isn't a public API, so check it still works after you upgrade Blume.

npm install --save-dev playwright@1.63.0
npx playwright install chromium
// Runs every query in search-queries.json through a Blume site's search
// dialog and records where the first correct page ranks.
// Usage: node search-eval.mjs <site-url> <candidate-name>
import { existsSync } from "node:fs";
import { appendFile, readFile } from "node:fs/promises";
import { chromium } from "playwright";

const [site, candidate] = process.argv.slice(2);
const queries = JSON.parse(await readFile("search-queries.json", "utf8"));
const date = new Date().toLocaleDateString("en-CA");
const pathOf = (href) => {
  const path = new URL(href).pathname;
  return path.length > 1 && path.endsWith("/") ? path.slice(0, -1) : path;
};

if (!existsSync("results.csv")) {
  await appendFile("results.csv", "date,candidate,query,kind,rank,results\n");
}

const browser = await chromium.launch();
const page = await browser.newPage();
await page.goto(site);
await page.keyboard.press("/");
const input = page.locator("[data-blume-search-input]");
await input.waitFor();

let top1 = 0;
let top3 = 0;
for (const { query, kind, expect } of queries) {
  await input.fill(query);
  // Settled once a row renders, or a message other than the loading "…".
  await page.waitForFunction(() => {
    const row = document.querySelector("#blume-search-listbox [role=option]");
    const note = document.querySelector("[data-blume-search-message]");
    return row || (note && !note.hidden && note.textContent !== "…");
  });
  const urls = await page.$$eval("#blume-search-listbox a[role=option]", (links) =>
    links.map((link) => link.href)
  );
  const rank = urls.map(pathOf).findIndex((url) => expect.includes(url)) + 1;
  if (rank === 1) top1 += 1;
  if (rank > 0 && rank <= 3) top3 += 1;
  const cells = [date, candidate, JSON.stringify(query), kind, rank || "miss", urls.length];
  await appendFile("results.csv", `${cells.join(",")}\n`);
}
await browser.close();
console.log(`${candidate}: top 1 ${top1}/${queries.length}, top 3 ${top3}/${queries.length}`);

Build each candidate, changing only search in blume.config.ts, and serve it with blume preview, which listens on port 4321. Start with the default, since it's your baseline:

npx blume build
npx blume preview
# in a second terminal:
node search-eval.mjs http://localhost:4321 orama

Stop the preview before each new build. For Pagefind, set search: pagefind() and repeat. Always test a build, since Pagefind's index doesn't exist in blume dev. Pagefind's result links end in a slash, like /docs/deployment/, which blume preview answers with Not Found. The script strips the slash before it compares, so the ranks still count; the Pagefind guide covers what that means on your host.

A hosted keyword candidate

Pick one provider to stand for the category. With Algolia, create an index, put the search-only key in the config, and put the admin key in .env.local, which Blume loads for every command:

import { defineConfig } from "blume";
import { algolia } from "blume/search";

export default defineConfig({
  search: algolia({
    appId: "YOUR_APP_ID",
    apiKey: "YOUR_SEARCH_ONLY_KEY",
    indexName: "docs_eval",
  }),
});
npm install algoliasearch
echo "ALGOLIA_ADMIN_API_KEY=your-admin-key" >> .env.local
npx blume build
npx blume preview

The build log should say Synced 100 record(s) to algolia, with your page count. If it says Search sync skipped: instead, nothing was uploaded and the results don't count.

A semantic candidate

Mixedbread's queries go through a /api/search route on your server, so this candidate needs a host adapter. Use node(), since blume preview can serve its build locally:

import { defineConfig } from "blume";
import { node } from "blume/deploy";
import { mixedbread } from "blume/search";

export default defineConfig({
  deployment: node(),
  search: mixedbread({ storeId: "acme-docs" }),
});

Blume doesn't upload anything to Mixedbread. You fill the store with Mixedbread's CLI, which reads MXBAI_API_KEY, while Blume's route reads MIXEDBREAD_API_KEY:

npm install @mixedbread/sdk
npm install --save-dev @mixedbread/cli@2.4.0
export MXBAI_API_KEY=your-mixedbread-key
npx mxbai store create acme-docs
npx mxbai store sync acme-docs "docs/**" --yes
echo "MIXEDBREAD_API_KEY=your-mixedbread-key" >> .env.local
npx blume build
npx blume preview

Compare more than rank

The top 1 and top 3 counts are the headline, but read the misses by kind: an engine that only misses typos has a different problem from one that misses every question. Then check what the counts hide.

  • Links. Blume's Mixedbread route takes each result's link from a url in the metadata Mixedbread generated for the file, which neither mxbai store sync nor Blume sets. A result without one links back to the current page, and the script counts it as a miss. Check what your store returns with npx mxbai store search acme-docs "dark mode" --return-metadata --format json before you trust semantic results.
  • Scoping. Run the script again from a translated page and from a page in an older version, and check that results stay in that language and version where the table says they should. The "All languages" toggle shows on every adapter, but changes nothing on Pagefind, which keeps to the page's language, or Mixedbread, which searches every language.
  • Freshness. A local index is rebuilt with the site. A hosted index holds whatever the last successful upload sent, and a failed upload only warns. Typesense is the exception: Blume deletes the collection before it uploads, so searches fail while a sync runs, and a sync that fails partway leaves the collection missing or partial. The Mixedbread store changes only when someone runs mxbai store sync, so it needs its own CI step.

What we measured on Blume's docs

We ran the 12 queries on 2026-09-27 against Blume's own docs on the main branch: 100 English pages, docs plus changelog. Orama, FlexSearch, and Typesense ran through Blume's own browser search clients, with Typesense 30.2 running locally and filled by Blume's upload, and Pagefind 1.5.2 indexed the pages as the dev server rendered them. The script above gave the same Orama row through the dialog. Mixedbread wasn't measured: it needs an account.

EngineTop 1Top 3Missed
Orama, the default456
FlexSearch556
Pagefind883
Typesense236
  • Questions were the weak spot everywhere: no engine found the deployment page for can I host the docs on my own server.
  • Only Pagefind found includeCodeBlocks, which appears only inside a code block, and only Pagefind put the right page first for both typos. Orama found nothing for either.
  • Adding chatbot to the assistant page's search.keywords moved it to first on Orama.
  • Typesense's three fifth-place results come from Blume's sort: relevance in ten bands, then search.boost. With every boost at 1, the llms.txt page landed behind four weaker matches. Sorted by relevance alone, the same index put 6 of 12 first.
  • The Orama index ships whole to each reader on first search, every language in one file: 3.0 MB for 300 pages in five languages, about 870 KB with gzip. Pagefind's first query loaded about 210 KB, plus its 45 KB script.
  • With English as the default locale, the Japanese and Hindi pages returned nothing on Orama, even for their own titles. German worked.
  • As Algolia records, 23 of the 100 pages are over 10 KB. Algolia caps record size by plan, from 10 KB to 100 KB (checked 2026-09-27), and Blume uploads one record per page.

These are observations from one site and 12 queries, not a ranking of engines. Pagefind's lead here came mostly from a code term and two typos, and one keyword closed one of Orama's misses. Run your own set before you switch.

Which search fits your docs

  • Most docs: keep Orama. Fix misses with search.keywords, search.boost, and includeCodeBlocks first. Keywords and boosts also rank pages for the assistant and the MCP server.
  • Large or many-language static sites, or terms that live in code: Pagefind. You give up search in blume dev and version scoping. See Add Pagefind search.
  • Versioned docs: Orama, FlexSearch, Algolia, or Typesense. The others search every version at once.
  • Pages in a non-Latin script: with a non-Latin default language, Orama or Pagefind, not FlexSearch. With translations in another script than the default, like Japanese under English, Orama drops their words: use Pagefind, or Algolia with the index's languages set. The assistant can't search those words either way. See Make docs search work in Chinese, Japanese and Korean.
  • An index other tools also query, or ranking rules you manage at the provider: hosted keyword. Algolia keeps its index settings across uploads. Blume deletes and recreates the Typesense collection on every build, so settings tuned there need reapplying and searches fail while the sync runs (see Self-host documentation search with Typesense). Make sure someone sees the build's upload warning.
  • Readers ask in their own words and every keyword engine misses: semantic search, or the assistant. Both need server output (see Static or server-rendered documentation). With the assistant on, the dialog already offers every query to it as an Assistant row.

The assistant and MCP server keep their own index

Your search adapter changes the dialog, nothing else. The assistant grounds each answer in pages it retrieves from an Orama index baked into the build, whatever the adapter, even with search: false. The MCP server's search_docs tool and the JSON API's search use the same kind of index. So switching to Mixedbread doesn't make the assistant semantic, and a better dialog score says nothing about its answers. Only the inkeep() assistant adapter brings its own retrieval. To test answers instead of result lists, see Test whether your documentation can answer user questions.

Troubleshooting

Every query misses on a hosted candidate

Open the dialog yourself. "Something went wrong. Please try again." means the request to the provider failed: check the adapter's public key and host or endpoint. Mixedbread fails differently: when /api/search errors, the dialog shows no results instead of an error. Read the terminal running blume preview, and check that MIXEDBREAD_API_KEY is set; a build without it warns with BLUME_MISSING_SECRET.

The hosted index is missing new pages

Look for Search sync skipped: in the build log. It names the cause: an unset admin key, an oramaCloud() without indexId, or the provider rejecting the upload, such as an Algolia record over your plan's size limit. The build still succeeds, so the dialog keeps serving whatever the provider already holds. On Typesense, that can be a missing or partial collection, so build again once the cause is fixed.

The build fails with BLUME_SERVER_FEATURE_REQUIRED

mixedbread() on a static build fails with Search (mixedbread) requires server output. Set a host adapter from blume/deploy, like node().

The build fails with BLUME_DEPENDENCY_MISSING

The FlexSearch, Algolia, Typesense, Orama Cloud, and Mixedbread SDKs don't ship with Blume. Install the package the message names, like npm install algoliasearch.

Next step

Find your first queries

Pull the searches that found nothing from your analytics, then run them through your current search before you change anything.

Read the search analytics guide

A step here not working for you? Report a broken step.

Keep going.More guides.

Upgrade your docs with Blume.

Install today and ship a production-grade docs site in minutes. Free and open source, forever.

npx blume init