Skip to content
Blume
Esc
↑↓navigate↵open⌘Jpreview
Guides

Search

Make docs search work in Chinese, Japanese and Korean

A four-language test corpus, a query list to check it against, and a search setup that finds Chinese, Japanese and Korean words, with results for each engine.

By 11 min read

Blume's default search engine, Orama, reads every page with one tokenizer, and it picks that tokenizer from i18n.defaultLocale. When the default is English or another Latin-script language, Chinese, Japanese and Korean text is dropped when the index is built. Only the Latin words and numbers in those pages survive, so a Japanese page turns up for RATE_LIMITED but never for 再試行. Make Chinese, Japanese or Korean the default and Orama switches to a word-segmenting tokenizer. Keep English as the default, and Pagefind is the adapter that gives each language its own segmented index.

By the end of this guide you have a small test site in four languages, a list of queries with the pages each should find, and a search setup that passes it. You also get the results we recorded for each engine on that corpus: exact terms, parts of words, and queries that mix code with text.

If your docs are in one Latin-script language, the default already works. The results below hold for the listed queries on the listed runtimes, not for every sentence in every language. Blume doesn't convert between Simplified and Traditional characters, match kana readings to kanji, or strip Korean particles. If you need that, use a hosted engine with its own language processing, and test it the same way.

Why CJK words don't match

Orama's standard tokenizer splits text on anything that isn't a basic Latin letter, a digit, or one of a few accented vowels. That's not only a spacing problem: every Han, kana and Hangul character counts as a separator, so a page in those scripts indexes none of its own words.

When the default locale resolves to a non-Latin script, Blume swaps in a tokenizer built on Intl.Segmenter, the word segmenter built into browsers and Node.js. For Chinese and Japanese defaults it also indexes Han, hiragana and katakana as overlapping pairs of characters, so 署名検証 is stored as 署名, 名検 and 検証. A query first looks for pages that carry all of its pairs, which keeps a page that mentions 署名 and 検証 in separate sentences out of the results for the compound. Korean indexes keep the segmenter's words.

The adapters don't all do this the same way:

AdapterCJK textLanguages
Orama (default)Segmented only when the default locale is non-Latin, with character pairs for Chinese and JapaneseOne tokenizer for every language, from the default locale
FlexSearchNot segmented: a run of CJK text matches only from its startOne index, filtered by locale in the browser
PagefindSegmented for Chinese, Japanese and Korean by the extended binary its npm package installsOne index per <html lang>
AlgoliaSegmented once you set the index's languages yourselfOne index, filtered by a locale facet
TypesenseBlume creates the collection without a field locale, so Typesense's default tokenizer, made for languages that put spaces between words, applies. Blume recreates the collection on every build, so a locale you add by hand doesn't last.One collection, filtered by a locale facet

Orama Cloud and Mixedbread do their language processing on their side, and Blume doesn't configure it, so test them with your corpus before you rely on them.

Build a test corpus

You can't judge search on your real docs, because you don't know which pages each query should find. Build a small site where you do. This script creates three Acme pages in English, Japanese, Simplified Chinese and Korean. Each page mixes prose with code terms, and the API keys page mentions 署名 (signature) and 検証 (verification) in separate sentences, so it shows whether a compound query matches loosely.

mkdir -p docs/ja docs/zh docs/ko

cat > docs/index.md <<'EOF'
---
title: Acme docs
description: Guides for the Acme Messages API.
---

Start with [webhooks](/webhooks), [rate limits](/rate-limits), or [API keys](/api-keys).
EOF

cat > docs/webhooks.md <<'EOF'
---
title: Verify webhook signatures
description: Confirm that a webhook came from Acme.
---

Every webhook request carries an `Acme-Signature` header. Pass your `signingSecret` to `acme.webhooks.verify()` to verify the signature. Return `401` when verification fails.
EOF

cat > docs/rate-limits.md <<'EOF'
---
title: Rate limits
description: Request limits and how to retry.
---

Over the limit, the API returns `429` with a `RATE_LIMITED` error. Wait for the number of seconds in the `Retry-After` header, then retry.
EOF

cat > docs/api-keys.md <<'EOF'
---
title: Manage API keys
description: Create, store, and rotate API keys.
---

Create API keys on the settings page of the dashboard. Requests are signed with your API key, and the key is checked on every request.
EOF

cat > docs/ja/webhooks.md <<'EOF'
---
title: Webhook の署名を検証する
description: 受信した Webhook が Acme から送信されたことを確認します。
---

すべての Webhook リクエストには `Acme-Signature` ヘッダーが付きます。`signingSecret` を渡して `acme.webhooks.verify()` を呼び出し、署名検証を行います。検証に失敗した場合は `401` を返してください。
EOF

cat > docs/ja/rate-limits.md <<'EOF'
---
title: レート制限
description: リクエスト数の上限と再試行の方法。
---

上限を超えると、API は `RATE_LIMITED` エラーとともに `429` を返します。`Retry-After` ヘッダーに示された秒数だけ待ってから再試行してください。
EOF

cat > docs/ja/api-keys.md <<'EOF'
---
title: API キーを管理する
description: API キーの作成、保管、ローテーション。
---

API キーはダッシュボードの設定画面で作成します。リクエストの署名には API キーを使用し、キーはリクエストのたびに検証されます。
EOF

cat > docs/zh/webhooks.md <<'EOF'
---
title: 验证 Webhook 签名
description: 确认 Webhook 来自 Acme。
---

每个 Webhook 请求都带有 `Acme-Signature` 请求头。调用 `acme.webhooks.verify()` 并传入 `signingSecret`,即可完成签名验证。验证失败时,请返回 `401`。
EOF

cat > docs/zh/rate-limits.md <<'EOF'
---
title: 速率限制
description: 请求数量上限以及重试方法。
---

超出上限后,API 会返回 `429` 和 `RATE_LIMITED` 错误。请等待 `Retry-After` 请求头中给出的秒数,然后重试。
EOF

cat > docs/zh/api-keys.md <<'EOF'
---
title: 管理 API 密钥
description: 创建、保存和轮换 API 密钥。
---

在控制台的设置页面创建 API 密钥。请求使用 API 密钥签名,每次请求都会验证密钥。
EOF

cat > docs/ko/webhooks.md <<'EOF'
---
title: Webhook 서명 검증하기
description: Webhook이 Acme에서 보낸 것인지 확인합니다.
---

모든 Webhook 요청에는 `Acme-Signature` 헤더가 포함됩니다. `signingSecret`을 전달해 `acme.webhooks.verify()`를 호출하면 서명을 검증할 수 있습니다. 검증에 실패하면 `401`을 반환하세요.
EOF

cat > docs/ko/rate-limits.md <<'EOF'
---
title: 속도 제한
description: 요청 한도와 재시도 방법.
---

한도를 초과하면 API가 `RATE_LIMITED` 오류와 함께 `429`를 반환합니다. `Retry-After` 헤더에 표시된 초만큼 기다린 뒤 재시도하세요.
EOF

cat > docs/ko/api-keys.md <<'EOF'
---
title: API 키 관리하기
description: API 키 생성, 보관, 교체.
---

API 키는 대시보드의 설정 페이지에서 만듭니다. 요청은 API 키로 서명되며, 키는 요청할 때마다 검증됩니다.
EOF

Run it in an empty folder with sh create-corpus.sh, then install Blume there with npm install blume.

Write the queries down first

Decide what each query should find before you look at any results. Keep the list in a file next to the corpus:

language	query	tests	should find
ja	署名検証	compound term	webhooks
ja	検証	part of a compound	webhooks, api-keys
ja	再試行	word mid-sentence	rate-limits
ja	RATE_LIMITED	code term	rate-limits
ja	Webhook 署名	code and text	webhooks
ja	けんしょう	reading in kana	nothing
zh	签名验证	compound term	webhooks
zh	重试	word mid-sentence	rate-limits
zh	簽名驗證	Traditional characters	webhooks
ko	서명	word stem	webhooks, api-keys
ko	서명을	stem with a particle	webhooks
ko	속도제한	space left out	rate-limits

Ask a native speaker of each language to review the pages and the list, and to add the queries your readers type: product names in katakana, terms your team writes both with and without spaces, the Traditional forms your Taiwan or Hong Kong readers use. A test list written by someone who doesn't read the language tends to test only what the tokenizer already handles.

Reproduce the failure

Add a config with English as the default and the three translations, and leave search out, so the default Orama adapter runs:

import { defineConfig } from "blume";

export default defineConfig({
  title: "Acme Docs",
  i18n: {
    defaultLocale: "en",
    locales: [
      { code: "en", label: "English" },
      { code: "ja", label: "日本語" },
      { code: "zh", label: "简体中文" },
      { code: "ko", label: "한국어" },
    ],
  },
});

Start npx blume dev, open /ja/rate-limits, and open search with ⌘K (or Ctrl K). On a page in a translation, the dialog searches that language only. Type 再試行: nothing. Type RATE_LIMITED: the Japanese page comes up. That's the signature of this failure. The code terms in a page are Latin, so they survive, while every Japanese word is gone. Webhook 署名 also finds the Japanese webhooks page, but only because Webhook matched: the tokenizer dropped 署名 from the query too.

Pick the path for your default language

Chinese, Japanese or Korean is your main language

Make it the default. Orama then segments the whole index, and English pages stay searchable next to it, because Latin words come through segmentation intact:

import { defineConfig } from "blume";

export default defineConfig({
  title: "Acme Docs",
  i18n: {
    defaultLocale: "ja",
    locales: [
      { code: "ja", label: "日本語" },
      { code: "en", label: "English" },
    ],
  },
});

The default locale's pages live at the content root, and every other locale in a folder named by its code, so the Japanese pages move up to docs/ and the English ones into docs/en/. That changes URLs on a live site, so add redirects for the pages that move, as in the URL migration guide. The same default serves the assistant and the MCP server's search_docs, which use this index whatever your search setting is.

It's the script that decides, not the language name: zh, zh-TW and ja all get character pairs. A ko default segments into words instead, so on a Korean-default site with Chinese or Japanese translations, a compound like 署名検証 also matches pages that contain its parts separately.

English stays the default

Orama can't search the translations then, so switch the dialog to Pagefind. It reads the lang attribute Blume puts on each page's <html>, builds a separate index per language, and segments Chinese, Japanese and Korean at build time:

import { defineConfig } from "blume";
import { pagefind } from "blume/search";

export default defineConfig({
  title: "Acme Docs",
  i18n: {
    defaultLocale: "en",
    locales: [
      { code: "en", label: "English" },
      { code: "ja", label: "日本語" },
      { code: "zh", label: "简体中文" },
      { code: "ko", label: "한국어" },
    ],
  },
  search: pagefind(),
});

Pagefind only runs during a build, so build and preview the site:

npx blume build
ls dist/pagefind/*.pf_meta
npx blume preview

The build log says how many pages it indexed, and dist/pagefind holds one .pf_meta file per language, with the language code in its name (pagefind.ja_…). A search on a Japanese page loads the Japanese index only, and the dialog's "All languages" toggle doesn't change that. Blume loads Pagefind once per full page load, and Pagefind picks the language as it loads, so after you switch language from the language menu, reload the page before you test. Pagefind has trade-offs of its own: no search in blume dev, no docs-version scoping, no results from custom pages, and result links that blume preview answers with Not Found. The Pagefind guide covers them.

Hosted search on Algolia

Algolia segments Chinese, Japanese and Korean with a dictionary once the index knows its languages. Algolia's docs say to set both indexLanguages and queryLanguages, and for CJK the first query language must be zh, ja, or ko. Blume's build sync doesn't set them, but it keeps the settings an index already has, so set them once:

import { algoliasearch } from "algoliasearch";

const indexName = "docs";
const client = algoliasearch(
  process.env.ALGOLIA_APP_ID,
  process.env.ALGOLIA_ADMIN_API_KEY
);

const { taskID } = await client.setSettings({
  indexName,
  indexSettings: {
    indexLanguages: ["ja", "en"],
    queryLanguages: ["ja", "en"],
  },
});
await client.waitForTask({ indexName, taskID });

Install the client with npm install algoliasearch@5.59.0, set ALGOLIA_APP_ID and ALGOLIA_ADMIN_API_KEY, run the script with node set-languages.mjs, and then run blume build again so the next sync uploads the pages into the configured index. Every language shares that one index, so with more than one CJK language, put your main one first and check the others against your query list. We didn't run Algolia for this guide.

Compare the results

We ran the query list against the corpus text on each engine. Orama ran Blume's own index code on Node.js 24.14.1 (ICU 78.2), Bun 1.4.2 (ICU 78.1), and headless Chromium 153 and WebKit 26.6 on macOS, with the same results on all four. Pagefind 1.5.2 ran its extended binary and was queried in Chromium from a page in each query's language. FlexSearch 0.8.212 ran with the options Blume gives it. Each query was scoped to its own language, the way the dialog scopes it on a translated page.

QueryOrama, English defaultOrama, Japanese defaultFlexSearchPagefind
署名検証nothingwebhookswebhookswebhooks, api-keys
検証nothingwebhooks, api-keyswebhookswebhooks, api-keys
再試行nothingrate-limitsnothingrate-limits
RATE_LIMITEDrate-limitsrate-limitsrate-limitsrate-limits
Webhook 署名webhooks, on Webhook alonewebhookswebhookswebhooks
けんしょうnothingnothingnothingall three pages
签名验证nothingwebhooksnothingwebhooks, api-keys
重试nothingrate-limitsnothingrate-limits
簽名驗證nothingnothingnothingnothing
서명nothingwebhooks, api-keyswebhooks, api-keyswebhooks, api-keys
서명을nothingwebhookswebhookswebhooks
속도제한nothingnothingnothingrate-limits

What the table shows, for these queries:

  • Orama with a Japanese default passed every Japanese and Chinese query except the Traditional one, and it was the only engine that kept the API keys page out of the compound results. A Chinese default gave the same results; a Korean default also returned the API keys page for 署名検証 and 签名验证.
  • Pagefind found every page it should for the words, and it was the only engine to match the Korean compound typed without its space. It splits a compound into words that can match anywhere on a page, but it ranked the webhooks page first, and the quoted query "署名検証" returned only that page.
  • FlexSearch found a Chinese or Japanese word only when it started a run of text between spaces or punctuation, so it missed words in the middle of a sentence. Don't use it for CJK content.
  • No engine converted Traditional characters to Simplified, or matched a kana reading to its kanji.

Troubleshooting

A translated page matches its code terms but not its words

Your default locale is Latin-script, and you're on Orama. Make the CJK language the default or switch to Pagefind, as above. Adding CJK terms to a page's search.keywords doesn't help, because keywords go through the same tokenizer.

Search says it's available in the production build

That's Pagefind in blume dev. The index only exists after blume build, so test with blume preview.

A query returns pages that don't contain it

Pagefind splits a query into words, and a page can match on part of them. In our run, けんしょう returned all three Japanese pages, none of which contains it; each excerpt highlighted only し. Wrap a term in double quotes for an exact match. Orama loosens too: when no page carries all of a query's character pairs, it returns pages that carry some, so a whole sentence still finds its closest pages.

A Traditional Chinese query misses Simplified pages

Neither engine converts between the two. Add the Traditional form to the page's frontmatter, which worked for both Orama with a Chinese or Japanese default and Pagefind in our run:

search:
  keywords: [簽名驗證, 驗證]

A single character finds nothing

Orama with a Chinese or Japanese default matches one character only where it starts a character pair, so a character at the end of a word can miss. Pagefind returned nothing for 証 in our run. Type two or more characters.

A page never appears in its language's results

Check that the page is really translated. An untranslated page is served as a copy of the default locale's page, and those copies are left out of the search index.

Results differ in another browser or on CI

Orama builds its index in each reader's browser from blume-search.json, and Pagefind segments queries in the browser too, both with the browser's own Intl.Segmenter. ICU versions differ between browsers and operating systems, and segmentation can change with them. Run your query list in the browsers your readers use, not only on the machine that passed it.

Next step

Track searches with no results

With an analytics adapter on, every query that finds nothing is recorded with results: 0, so the CJK terms your test list missed show up.

Read the search analytics guide

A step here not working for you? Report a broken step.

Keep going.More guides.

Upgrade your docs with Blume.

Install today and ship a production-grade docs site in minutes. Free and open source, forever.

npx blume init