Search
Make docs search work in Chinese, Japanese and Korean
A four-language test corpus, a query list to check it against, and a search setup that finds Chinese, Japanese and Korean words, with results for each engine.
By Hayden Bleasel11 min read

Blume's default search engine, Orama, reads every page with one tokenizer, and it picks that tokenizer from i18n.defaultLocale. When the default is English or another Latin-script language, Chinese, Japanese and Korean text is dropped when the index is built. Only the Latin words and numbers in those pages survive, so a Japanese page turns up for RATE_LIMITED but never for 再試行. Make Chinese, Japanese or Korean the default and Orama switches to a word-segmenting tokenizer. Keep English as the default, and Pagefind is the adapter that gives each language its own segmented index.
By the end of this guide you have a small test site in four languages, a list of queries with the pages each should find, and a search setup that passes it. You also get the results we recorded for each engine on that corpus: exact terms, parts of words, and queries that mix code with text.
If your docs are in one Latin-script language, the default already works. The results below hold for the listed queries on the listed runtimes, not for every sentence in every language. Blume doesn't convert between Simplified and Traditional characters, match kana readings to kanji, or strip Korean particles. If you need that, use a hosted engine with its own language processing, and test it the same way.
Why CJK words don't match
Orama's standard tokenizer splits text on anything that isn't a basic Latin letter, a digit, or one of a few accented vowels. That's not only a spacing problem: every Han, kana and Hangul character counts as a separator, so a page in those scripts indexes none of its own words.
When the default locale resolves to a non-Latin script, Blume swaps in a tokenizer built on Intl.Segmenter, the word segmenter built into browsers and Node.js. For Chinese and Japanese defaults it also indexes Han, hiragana and katakana as overlapping pairs of characters, so 署名検証 is stored as 署名, 名検 and 検証. A query first looks for pages that carry all of its pairs, which keeps a page that mentions 署名 and 検証 in separate sentences out of the results for the compound. Korean indexes keep the segmenter's words.
The adapters don't all do this the same way:
| Adapter | CJK text | Languages |
|---|---|---|
| Orama (default) | Segmented only when the default locale is non-Latin, with character pairs for Chinese and Japanese | One tokenizer for every language, from the default locale |
| FlexSearch | Not segmented: a run of CJK text matches only from its start | One index, filtered by locale in the browser |
| Pagefind | Segmented for Chinese, Japanese and Korean by the extended binary its npm package installs | One index per <html lang> |
| Algolia | Segmented once you set the index's languages yourself | One index, filtered by a locale facet |
| Typesense | Blume creates the collection without a field locale, so Typesense's default tokenizer, made for languages that put spaces between words, applies. Blume recreates the collection on every build, so a locale you add by hand doesn't last. | One collection, filtered by a locale facet |
Orama Cloud and Mixedbread do their language processing on their side, and Blume doesn't configure it, so test them with your corpus before you rely on them.
Build a test corpus
You can't judge search on your real docs, because you don't know which pages each query should find. Build a small site where you do. This script creates three Acme pages in English, Japanese, Simplified Chinese and Korean. Each page mixes prose with code terms, and the API keys page mentions 署名 (signature) and 検証 (verification) in separate sentences, so it shows whether a compound query matches loosely.
mkdir -p docs/ja docs/zh docs/ko
cat > docs/index.md <<'EOF'
---
title: Acme docs
description: Guides for the Acme Messages API.
---
Start with [webhooks](/webhooks), [rate limits](/rate-limits), or [API keys](/api-keys).
EOF
cat > docs/webhooks.md <<'EOF'
---
title: Verify webhook signatures
description: Confirm that a webhook came from Acme.
---
Every webhook request carries an `Acme-Signature` header. Pass your `signingSecret` to `acme.webhooks.verify()` to verify the signature. Return `401` when verification fails.
EOF
cat > docs/rate-limits.md <<'EOF'
---
title: Rate limits
description: Request limits and how to retry.
---
Over the limit, the API returns `429` with a `RATE_LIMITED` error. Wait for the number of seconds in the `Retry-After` header, then retry.
EOF
cat > docs/api-keys.md <<'EOF'
---
title: Manage API keys
description: Create, store, and rotate API keys.
---
Create API keys on the settings page of the dashboard. Requests are signed with your API key, and the key is checked on every request.
EOF
cat > docs/ja/webhooks.md <<'EOF'
---
title: Webhook の署名を検証する
description: 受信した Webhook が Acme から送信されたことを確認します。
---
すべての Webhook リクエストには `Acme-Signature` ヘッダーが付きます。`signingSecret` を渡して `acme.webhooks.verify()` を呼び出し、署名検証を行います。検証に失敗した場合は `401` を返してください。
EOF
cat > docs/ja/rate-limits.md <<'EOF'
---
title: レート制限
description: リクエスト数の上限と再試行の方法。
---
上限を超えると、API は `RATE_LIMITED` エラーとともに `429` を返します。`Retry-After` ヘッダーに示された秒数だけ待ってから再試行してください。
EOF
cat > docs/ja/api-keys.md <<'EOF'
---
title: API キーを管理する
description: API キーの作成、保管、ローテーション。
---
API キーはダッシュボードの設定画面で作成します。リクエストの署名には API キーを使用し、キーはリクエストのたびに検証されます。
EOF
cat > docs/zh/webhooks.md <<'EOF'
---
title: 验证 Webhook 签名
description: 确认 Webhook 来自 Acme。
---
每个 Webhook 请求都带有 `Acme-Signature` 请求头。调用 `acme.webhooks.verify()` 并传入 `signingSecret`,即可完成签名验证。验证失败时,请返回 `401`。
EOF
cat > docs/zh/rate-limits.md <<'EOF'
---
title: 速率限制
description: 请求数量上限以及重试方法。
---
超出上限后,API 会返回 `429` 和 `RATE_LIMITED` 错误。请等待 `Retry-After` 请求头中给出的秒数,然后重试。
EOF
cat > docs/zh/api-keys.md <<'EOF'
---
title: 管理 API 密钥
description: 创建、保存和轮换 API 密钥。
---
在控制台的设置页面创建 API 密钥。请求使用 API 密钥签名,每次请求都会验证密钥。
EOF
cat > docs/ko/webhooks.md <<'EOF'
---
title: Webhook 서명 검증하기
description: Webhook이 Acme에서 보낸 것인지 확인합니다.
---
모든 Webhook 요청에는 `Acme-Signature` 헤더가 포함됩니다. `signingSecret`을 전달해 `acme.webhooks.verify()`를 호출하면 서명을 검증할 수 있습니다. 검증에 실패하면 `401`을 반환하세요.
EOF
cat > docs/ko/rate-limits.md <<'EOF'
---
title: 속도 제한
description: 요청 한도와 재시도 방법.
---
한도를 초과하면 API가 `RATE_LIMITED` 오류와 함께 `429`를 반환합니다. `Retry-After` 헤더에 표시된 초만큼 기다린 뒤 재시도하세요.
EOF
cat > docs/ko/api-keys.md <<'EOF'
---
title: API 키 관리하기
description: API 키 생성, 보관, 교체.
---
API 키는 대시보드의 설정 페이지에서 만듭니다. 요청은 API 키로 서명되며, 키는 요청할 때마다 검증됩니다.
EOFRun it in an empty folder with sh create-corpus.sh, then install Blume there with npm install blume.
Write the queries down first
Decide what each query should find before you look at any results. Keep the list in a file next to the corpus:
language query tests should find
ja 署名検証 compound term webhooks
ja 検証 part of a compound webhooks, api-keys
ja 再試行 word mid-sentence rate-limits
ja RATE_LIMITED code term rate-limits
ja Webhook 署名 code and text webhooks
ja けんしょう reading in kana nothing
zh 签名验证 compound term webhooks
zh 重试 word mid-sentence rate-limits
zh 簽名驗證 Traditional characters webhooks
ko 서명 word stem webhooks, api-keys
ko 서명을 stem with a particle webhooks
ko 속도제한 space left out rate-limitsAsk a native speaker of each language to review the pages and the list, and to add the queries your readers type: product names in katakana, terms your team writes both with and without spaces, the Traditional forms your Taiwan or Hong Kong readers use. A test list written by someone who doesn't read the language tends to test only what the tokenizer already handles.
Reproduce the failure
Add a config with English as the default and the three translations, and leave search out, so the default Orama adapter runs:
import { defineConfig } from "blume";
export default defineConfig({
title: "Acme Docs",
i18n: {
defaultLocale: "en",
locales: [
{ code: "en", label: "English" },
{ code: "ja", label: "日本語" },
{ code: "zh", label: "简体中文" },
{ code: "ko", label: "한국어" },
],
},
});Start npx blume dev, open /ja/rate-limits, and open search with ⌘K (or Ctrl K). On a page in a translation, the dialog searches that language only. Type 再試行: nothing. Type RATE_LIMITED: the Japanese page comes up. That's the signature of this failure. The code terms in a page are Latin, so they survive, while every Japanese word is gone. Webhook 署名 also finds the Japanese webhooks page, but only because Webhook matched: the tokenizer dropped 署名 from the query too.
Pick the path for your default language
Chinese, Japanese or Korean is your main language
Make it the default. Orama then segments the whole index, and English pages stay searchable next to it, because Latin words come through segmentation intact:
import { defineConfig } from "blume";
export default defineConfig({
title: "Acme Docs",
i18n: {
defaultLocale: "ja",
locales: [
{ code: "ja", label: "日本語" },
{ code: "en", label: "English" },
],
},
});The default locale's pages live at the content root, and every other locale in a folder named by its code, so the Japanese pages move up to docs/ and the English ones into docs/en/. That changes URLs on a live site, so add redirects for the pages that move, as in the URL migration guide. The same default serves the assistant and the MCP server's search_docs, which use this index whatever your search setting is.
It's the script that decides, not the language name: zh, zh-TW and ja all get character pairs. A ko default segments into words instead, so on a Korean-default site with Chinese or Japanese translations, a compound like 署名検証 also matches pages that contain its parts separately.
English stays the default
Orama can't search the translations then, so switch the dialog to Pagefind. It reads the lang attribute Blume puts on each page's <html>, builds a separate index per language, and segments Chinese, Japanese and Korean at build time:
import { defineConfig } from "blume";
import { pagefind } from "blume/search";
export default defineConfig({
title: "Acme Docs",
i18n: {
defaultLocale: "en",
locales: [
{ code: "en", label: "English" },
{ code: "ja", label: "日本語" },
{ code: "zh", label: "简体中文" },
{ code: "ko", label: "한국어" },
],
},
search: pagefind(),
});Pagefind only runs during a build, so build and preview the site:
npx blume build
ls dist/pagefind/*.pf_meta
npx blume previewThe build log says how many pages it indexed, and dist/pagefind holds one .pf_meta file per language, with the language code in its name (pagefind.ja_…). A search on a Japanese page loads the Japanese index only, and the dialog's "All languages" toggle doesn't change that. Blume loads Pagefind once per full page load, and Pagefind picks the language as it loads, so after you switch language from the language menu, reload the page before you test. Pagefind has trade-offs of its own: no search in blume dev, no docs-version scoping, no results from custom pages, and result links that blume preview answers with Not Found. The Pagefind guide covers them.
Hosted search on Algolia
Algolia segments Chinese, Japanese and Korean with a dictionary once the index knows its languages. Algolia's docs say to set both indexLanguages and queryLanguages, and for CJK the first query language must be zh, ja, or ko. Blume's build sync doesn't set them, but it keeps the settings an index already has, so set them once:
import { algoliasearch } from "algoliasearch";
const indexName = "docs";
const client = algoliasearch(
process.env.ALGOLIA_APP_ID,
process.env.ALGOLIA_ADMIN_API_KEY
);
const { taskID } = await client.setSettings({
indexName,
indexSettings: {
indexLanguages: ["ja", "en"],
queryLanguages: ["ja", "en"],
},
});
await client.waitForTask({ indexName, taskID });Install the client with npm install algoliasearch@5.59.0, set ALGOLIA_APP_ID and ALGOLIA_ADMIN_API_KEY, run the script with node set-languages.mjs, and then run blume build again so the next sync uploads the pages into the configured index. Every language shares that one index, so with more than one CJK language, put your main one first and check the others against your query list. We didn't run Algolia for this guide.
Compare the results
We ran the query list against the corpus text on each engine. Orama ran Blume's own index code on Node.js 24.14.1 (ICU 78.2), Bun 1.4.2 (ICU 78.1), and headless Chromium 153 and WebKit 26.6 on macOS, with the same results on all four. Pagefind 1.5.2 ran its extended binary and was queried in Chromium from a page in each query's language. FlexSearch 0.8.212 ran with the options Blume gives it. Each query was scoped to its own language, the way the dialog scopes it on a translated page.
| Query | Orama, English default | Orama, Japanese default | FlexSearch | Pagefind |
|---|---|---|---|---|
| 署名検証 | nothing | webhooks | webhooks | webhooks, api-keys |
| 検証 | nothing | webhooks, api-keys | webhooks | webhooks, api-keys |
| 再試行 | nothing | rate-limits | nothing | rate-limits |
RATE_LIMITED | rate-limits | rate-limits | rate-limits | rate-limits |
| Webhook 署名 | webhooks, on Webhook alone | webhooks | webhooks | webhooks |
| けんしょう | nothing | nothing | nothing | all three pages |
| 签名验证 | nothing | webhooks | nothing | webhooks, api-keys |
| 重试 | nothing | rate-limits | nothing | rate-limits |
| 簽名驗證 | nothing | nothing | nothing | nothing |
| 서명 | nothing | webhooks, api-keys | webhooks, api-keys | webhooks, api-keys |
| 서명을 | nothing | webhooks | webhooks | webhooks |
| 속도제한 | nothing | nothing | nothing | rate-limits |
What the table shows, for these queries:
- Orama with a Japanese default passed every Japanese and Chinese query except the Traditional one, and it was the only engine that kept the API keys page out of the compound results. A Chinese default gave the same results; a Korean default also returned the API keys page for 署名検証 and 签名验证.
- Pagefind found every page it should for the words, and it was the only engine to match the Korean compound typed without its space. It splits a compound into words that can match anywhere on a page, but it ranked the webhooks page first, and the quoted query
"署名検証"returned only that page. - FlexSearch found a Chinese or Japanese word only when it started a run of text between spaces or punctuation, so it missed words in the middle of a sentence. Don't use it for CJK content.
- No engine converted Traditional characters to Simplified, or matched a kana reading to its kanji.
Troubleshooting
A translated page matches its code terms but not its words
Your default locale is Latin-script, and you're on Orama. Make the CJK language the default or switch to Pagefind, as above. Adding CJK terms to a page's search.keywords doesn't help, because keywords go through the same tokenizer.
Search says it's available in the production build
That's Pagefind in blume dev. The index only exists after blume build, so test with blume preview.
A query returns pages that don't contain it
Pagefind splits a query into words, and a page can match on part of them. In our run, けんしょう returned all three Japanese pages, none of which contains it; each excerpt highlighted only し. Wrap a term in double quotes for an exact match. Orama loosens too: when no page carries all of a query's character pairs, it returns pages that carry some, so a whole sentence still finds its closest pages.
A Traditional Chinese query misses Simplified pages
Neither engine converts between the two. Add the Traditional form to the page's frontmatter, which worked for both Orama with a Chinese or Japanese default and Pagefind in our run:
search:
keywords: [簽名驗證, 驗證]A single character finds nothing
Orama with a Chinese or Japanese default matches one character only where it starts a character pair, so a character at the end of a word can miss. Pagefind returned nothing for 証 in our run. Type two or more characters.
A page never appears in its language's results
Check that the page is really translated. An untranslated page is served as a copy of the default locale's page, and those copies are left out of the search index.
Results differ in another browser or on CI
Orama builds its index in each reader's browser from blume-search.json, and Pagefind segments queries in the browser too, both with the browser's own Intl.Segmenter. ICU versions differ between browsers and operating systems, and segmentation can change with them. Run your query list in the browsers your readers use, not only on the machine that passed it.
Next step
Track searches with no results
With an analytics adapter on, every query that finds nothing is recorded with results: 0, so the CJK terms your test list missed show up.
Read the search analytics guideA step here not working for you? Report a broken step.