Agents
Add an AI assistant that answers from your documentation
An in-page assistant that answers from your docs and links its sources, a support handoff for what they don't cover, and an evaluation set that shows where it fails.
By Hayden Bleasel11 min read

To add an AI chatbot that answers from your documentation and links to its sources, turn on Blume's assistant, pick a model provider, and deploy with a host adapter so its server route can run. Each question retrieves the most relevant pages of your docs, the model answers from them, and it cites the pages it used as links in the answer.
By the end, the Acme Messages API docs from Generate API docs from an OpenAPI spec have an assistant with suggested questions and a handoff to support, plus a small evaluation set you can run locally and in production. You'll also see a real run, failures included.
Grounding makes wrong answers less likely, not impossible. If you only need readers to find the right page, the built-in search already does that with no server and no model bill. If you want coding agents in an editor to read your docs, an MCP server is the better fit.
How the assistant finds an answer
At build time, Blume takes a snapshot of every page that's in search, as Markdown, even when search itself is off. A page with search.exclude: true in its frontmatter, or one hidden from the sidebar, stays out by default. ai.exclude doesn't keep a page from the assistant: it's for llms.txt.
For each question, the server route:
- Runs a keyword search over the snapshot, kept to the language (and docs version) of the page the reader is on.
- Puts that page first, then up to six matches, each cut down to its most relevant sections: 2,000 characters a page and 10,000 in all.
- Tells the model to answer only from those excerpts, to say when something isn't covered, and to cite each page as a Markdown link.
- Lets the model call
search_docsandread_pagefor up to four rounds before it answers.
Two things follow. Retrieval matches words, so a question phrased in words your docs never use can pull the wrong pages. And the instructions are text the model follows, not a check: nothing compares the finished answer against the excerpts.
Turn on the assistant
The built-in assistant is a server route, POST /api/ask, so it needs server output. Name your host's adapter from blume/deploy in the same config. This guide uses Vercel and the Vercel AI Gateway, Blume's default provider, which needs no extra package:
import { defineConfig } from "blume";
import { gateway } from "blume/ai";
import { vercel } from "blume/deploy";
export default defineConfig({
// ...the rest of your config
ai: {
assistant: {
enabled: true,
provider: gateway({ model: "openai/gpt-5.5" }),
suggestions: [
{ label: "How do I authenticate requests?", icon: "key-round" },
{ label: "How do I send an SMS?", icon: "message-square" },
{ label: "How do I check whether a message was delivered?", icon: "circle-check" },
],
instructions:
"You are the Acme Messages docs assistant. Keep answers short, and show the request when one helps. For billing or account questions, say the docs don't cover them and suggest contacting support.",
support: "mailto:support@acme.example",
},
},
deployment: vercel(),
});providerpicks who answers. This is the default, written out so you can see where to change the model. To call a provider like Anthropic directly, use its adapter.suggestionsare the questions a reader can click before typing. Take them from your support inbox: they're also your first test cases.instructionsare added after Blume's own instructions, never in place of them, so the rule to cite pages stays in force. They shape tone and scope. They can't make an answer true.supportadds a Contact support link under the conversation. Amailto:address opens an email holding the latest turns (up to 1,500 characters) and the conversation's id. A URL or a path on your site gets the id as athreadquery parameter instead.
Keep the key on the server
Create a key on the AI Gateway's API Keys page in your Vercel dashboard. For local development, put it in .env.local at the project root, which blume dev and blume build load:
AI_GATEWAY_API_KEY=your-gateway-keyThe .gitignore that blume init writes covers node_modules/, .blume/, and dist/, but not env files, so add it:
echo ".env.local" >> .gitignoreThe config holds only the variable's name. The generated route reads the value on the server at request time, so the key never lands in a built file or reaches the browser. An adapter's headers are written into the route as-is, so keep secrets out of those.
Until the key is set, blume dev and blume build warn with BLUME_MISSING_SECRET, and the route answers 503 with a message naming the variable, which the panel shows as is.
Try it locally
npx blume devOpen http://localhost:4321, then open the assistant from its button in the header, or with ⌘ I (Ctrl I on Windows and Linux). Click each suggestion and check three things:
- It cites pages. Sources show up as small pill-shaped links in the answer. An answer with none is worth a second look.
- The pages say what the answer says. A citation is a link the model wrote, not proof. Click each one.
- The current page counts. Go to
/api/messages/get-messageand ask "What does a 404 mean here?" The page you're on goes into the context first, so the answer should come from it.
Then ask something the docs don't cover, like "How do I get a refund for unused credits?" You want a plain "the docs don't cover this", not a guess. Click Contact support: your mail client should open a draft to support@acme.example with the conversation in it. The link shows under every conversation once an answer finishes, not only after a failed one.
Write a small evaluation set
Clicking around doesn't scale, and the same question can come back differently each time, so write the questions down. Each entry lists pages a good answer cites (cites, any one will do), words it must contain (mentions), and optionally the page it's asked from (page). Mark questions the docs shouldn't answer as review: only a person can judge a refusal.
[
{
"id": "authenticate",
"question": "How do I authenticate requests to the Acme API?",
"cites": ["/api/authentication"],
"mentions": ["Bearer"]
},
{
"id": "send-sms",
"question": "How do I send an SMS?",
"cites": ["/api/messages/send-message"],
"mentions": ["/messages"]
},
{
"id": "delivery-status",
"question": "How do I check whether a message was delivered?",
"cites": ["/api/messages/get-message"]
},
{
"id": "not-found-here",
"page": "/api/messages/get-message",
"question": "What does a 404 mean here?",
"cites": ["/api/messages/get-message"]
},
{
"id": "template-ids",
"question": "Where do I find my template IDs?",
"cites": ["/api/templates/list-templates"]
},
{ "id": "whatsapp", "question": "Can I send WhatsApp messages?", "review": true },
{ "id": "refund", "question": "How do I get a refund for unused credits?", "review": true }
]This script sends each question the way the panel does, reads the streamed answer, and checks it. It needs Node.js 22 or later and nothing else:
import { readFile } from "node:fs/promises";
const site = new URL(process.env.DOCS_URL ?? "http://localhost:4321");
const cases = JSON.parse(await readFile(process.argv[2] ?? "ask-evals.json", "utf8"));
const bypass = process.env.VERCEL_AUTOMATION_BYPASS_SECRET;
// The pages an answer links to on this site, without trailing slashes.
const citedRoutes = (answer) => {
const routes = new Set();
for (const [, href] of answer.matchAll(/\]\(([^)\s]+)\)/g)) {
const url = new URL(href, site);
if (url.origin === site.origin) {
routes.add(url.pathname.length > 1 ? url.pathname.replace(/\/$/, "") : "/");
}
}
return [...routes];
};
let failures = 0;
for (const test of cases) {
const response = await fetch(new URL("api/ask", `${site.href.replace(/\/$/, "")}/`), {
body: JSON.stringify({
messages: [{ content: test.question, role: "user" }],
page: { path: test.page ?? "/" },
}),
headers: {
"content-type": "application/json",
...(bypass && { "x-vercel-protection-bypass": bypass }),
},
method: "POST",
});
const answer = (await response.text()).trim();
const cited = citedRoutes(answer);
const problems = [];
if (!response.ok) {
problems.push(`HTTP ${response.status}`);
} else if (answer === "") {
problems.push("empty answer: check the server log");
} else {
if (test.cites && !test.cites.some((route) => cited.includes(route))) {
problems.push(`expected a link to ${test.cites.join(" or ")}`);
}
for (const phrase of test.mentions ?? []) {
if (!answer.toLowerCase().includes(phrase.toLowerCase())) {
problems.push(`never says "${phrase}"`);
}
}
}
const verdict = problems.length > 0 ? "FAIL" : test.review ? "REVIEW" : "PASS";
failures += verdict === "FAIL" ? 1 : 0;
console.log(`${verdict.padEnd(6)} ${test.id} cited: ${cited.join(", ") || "nothing"}`);
for (const problem of problems) {
console.log(` ${problem}`);
}
if (verdict !== "PASS") {
console.log(` answer: ${answer.replace(/\s+/g, " ").slice(0, 300)}`);
}
}
console.log(`\n${failures} of ${cases.length} failed`);
process.exitCode = failures > 0 ? 1 : 0;With blume dev running, in a second terminal:
node ask-eval.mjsEach line reads PASS, FAIL with the reason, or REVIEW with the answer, and any failure exits non-zero. A pass means "cited the right page and used the right words", not "correct": the checks match strings, so read the answers too.
What we saw on one site
In September 2026 we ran eight questions in this format against this site's own assistant on a local blume dev: the openai() adapter with the default retrieval and docs tools, over docs in five languages. That was three full runs, then two shorter ones. It's one site, one model, and a handful of runs, so read the table as the kinds of failure to look for, not as results to expect on yours.
| Question | What happened |
|---|---|
Suggested questions, the anthropic() key variable, deploying to Vercel | Passed every run, each citing the page that answers it. |
| "What is the default limit here?", asked from Rate limiting | Passed every run: it answered 30 and cited that page. |
| "How do I keep a page out of search results?" | Right answer every time, but it cited Search where our test expected Frontmatter. Both pages say it, so the test was wrong. We let it accept either. |
| "Does Blume support SAML single sign-on for readers?" | Asked from a docs page (once reworded), it passed both times, citing Deployment, whose Private docs section says Blume has no sign-in of its own. Asked from the home page, it said it didn't know all three times. |
| "How do I stop scripts from spamming the chat box?" | From a docs page, it mentioned the captcha option and cited the right pages twice. From the home page, it once suggested an unrelated setting and once said the docs don't cover it. Two runs came back empty. |
| "How do I turn on voice mode for the assistant?" | No such feature exists. Every answer said so and pointed to narration as a separate feature, which is what we wanted. |
In these runs, the difference was where the question was asked. This site's home page is a custom page, not a docs page. Retrieval and the model's own searches keep to the language of the reader's page, so from a page outside the docs they searched all five languages: for the single sign-on question, five of the six pages retrieved up front were translations. The fix was to send a docs page as page, as the panel does when a reader asks from one.
The empty answers were 200 responses with no text, which is how a provider error looks once an answer has started streaming. The last run also hit the rate limit and got a 429. Run the set more than once before you trust a pass or a fail.
Check the docs themselves with blume eval
ask-eval.mjs tests the assistant. blume eval tests the docs: Codex or Claude Code answers each question using only your pages, and a second session grades the answer against facts you list. It never calls your assistant. Use it to find what the docs don't say, and the script to see how the assistant handles what they do.
questions:
- id: authenticate
question: How do I authenticate requests to the Acme API?
expected:
- Send an API key as a bearer token in the Authorization header
routes: /api/authentication
- id: delivery-status
question: How do I check whether a message was delivered?
expected:
- Call Get a message with the ID returned by Send a message
routes: /api/messages/get-messagenpx blume eval --agent claudeIt uses your installed agent CLI and its account, not your gateway key. Test whether your documentation can answer user questions covers running it in CI.
Deploy and rerun in production
Add AI_GATEWAY_API_KEY to your Vercel project's environment variables for Production and Preview. (On Vercel, Blume also accepts the deployment's OIDC token instead.) Commit, push, and import the repository in Vercel with Node.js 22 or later, then point the script at it:
DOCS_URL=https://docs.acme.example node ask-eval.mjsPreviews behind Vercel Authentication return a login page instead of an answer. Create a Protection Bypass for Automation secret in the project settings and export it as VERCEL_AUTOMATION_BYPASS_SECRET; the script sends it as the x-vercel-protection-bypass header.
The snapshot is part of the build, so an edit to the docs reaches the assistant on the next deploy. If you later add a bot check, the script can't pass it and gets 403, so run it locally or on a deployment without one.
What it costs to run
Blume is free; the assistant bills in two places. Your model provider charges for every question: the excerpts and conversation go in, the answer comes out, and each round of searching or reading before the answer is another call. Your host charges for the server function that streams each answer. Closing the panel mid-answer stops the model call.
Rate limiting is on by default, at 30 questions per reader every 10 minutes, which also caps an eval run from one machine. On serverless hosts the default count is per instance: see Rate-limit a public documentation AI assistant to make it exact. The AI Gateway can also put a budget on a key. The pricing page estimates what answers cost in practice.
Troubleshooting
The build fails with BLUME_SERVER_FEATURE_REQUIRED
The assistant is on, but the build is static. Set deployment to a host adapter such as vercel(), or drop output: "static" from the one you have. To keep the site static, point endpoint at a backend you run instead.
The build stops with BLUME_DEPENDENCY_MISSING
Direct provider adapters need their AI SDK package, such as @ai-sdk/anthropic for anthropic(). The message includes the install command for your package manager. The gateway needs none.
The panel says "The assistant is not configured"
The route can't read the key and answered 503. Locally, check .env.local and restart blume dev. In production, set the variable in the host's environment and redeploy.
The panel says "Sorry, something went wrong"
Provider errors arrive after the answer has started streaming, so the reader gets this notice instead of an answer. The server log (the blume dev terminal, or your host's function logs) has the cause, after Assistant provider error:. The usual ones are a wrong model id, a revoked key, or the provider's own limits.
It says the docs don't cover something they do
Check that the page isn't set to search.exclude or hidden from the sidebar, and that the change is deployed. Otherwise it's usually wording: the reader's words aren't on the page. Use them in the text, or add them as search.keywords in the page's frontmatter, which retrieval weights like the title.
Answers cite a page in another language
On a multilingual site, retrieval keeps to the language of the page the question came from. Asked from a page outside the docs, like a custom landing page, it searches every language. Ask from a docs page, and set page on each eval question.
Next step
Test what your docs can answer
Have an agent draft an evals file from your existing pages, then run blume eval to find the questions your docs can't answer yet.
npx blume eval initA step here not working for you? Report a broken step.