Agents
Run a documentation assistant with Ollama
A docs assistant that answers from your pages with a model running in Ollama beside your docs server, sized for your hardware and measured against a hosted model.
By Hayden Bleasel10 min read

Blume's assistant can answer from your docs with a model you run in Ollama. Point the openai() adapter at Ollama's OpenAI-compatible Chat Completions endpoint, and Blume works as before: it finds the relevant pages, sends them with the question to your model, and streams back an answer that links its sources.
By the end, your docs server and Ollama run on one machine, Ollama is unreachable from outside it, the context window and retrieval fit your hardware, and you've asked a hosted model the same questions to see what you gave up.
A model that fits on one server usually answers less well than a hosted frontier model, and it moves the cost from a provider bill to hardware you keep running. If you don't need the model on your own machine, the hosted setup is less work. This guide reuses its questions for the comparison.
How the pieces connect
The assistant is a server route, POST /api/ask, on your docs site. For each question it searches a snapshot of your docs baked into the build, adds the best excerpts to the model's instructions, and calls the model. Browsers only talk to your docs server; only the route talks to Ollama.
So Ollama can stay where it starts, listening on 127.0.0.1:11434. Keep it there: Ollama ignores the API key on its local API, so anyone who can reach the port can run your model. If your docs are hosted elsewhere, see When the docs run somewhere else.
Local doesn't mean nothing leaves the machine. Models named :cloud or -cloud run on ollama.com, ollama pull downloads from Ollama's registry, and Blume features like analytics or a bot check talk to their own providers. If no traffic may leave, block it at the firewall and watch what gets refused.
Start Ollama with a model that calls tools
This guide pins Ollama 0.34.4 and uses qwen3:8b, a 5.2 GB download that Ollama lists with tool calling and thinking. On a Linux server, install that version with Ollama's script, which runs it as the ollama systemd service:
curl -fsSL https://ollama.com/install.sh | OLLAMA_VERSION=0.34.4 shOpen an override for the service with sudo systemctl edit ollama and add three settings:
[Service]
Environment="OLLAMA_CONTEXT_LENGTH=16384"
Environment="OLLAMA_KEEP_ALIVE=-1"
Environment="OLLAMA_NO_CLOUD=1"sudo systemctl daemon-reload
sudo systemctl restart ollama
ollama pull qwen3:8bThe context length makes room for Blume's excerpts (see Tune it for your hardware), the keep-alive stops Ollama unloading the model after five idle minutes, and the last line turns off cloud models. Leave OLLAMA_HOST unset so Ollama keeps listening on 127.0.0.1 only.
To try it on a Mac, run brew install ollama, quit the Ollama app if it's running (it holds the same port), and start the server with the same settings:
OLLAMA_CONTEXT_LENGTH=16384 OLLAMA_KEEP_ALIVE=-1 OLLAMA_NO_CLOUD=1 ollama serveWrite down what you're running. A tag like qwen3:8b can be updated upstream, so keep the ID column of ollama list next to the Ollama version. When answers change, you'll know whether the model did.
Test tool calling
Blume can give the model two tools, search_docs and read_page, to look past the excerpts it was handed. They're off by default for a self-hosted endpoint, since not every model can call tools. ollama show qwen3:8b should list tools under Capabilities. Then send a tool through the Chat Completions endpoint Blume will use:
curl -s http://127.0.0.1:11434/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3:8b",
"reasoning_effort": "none",
"messages": [{ "role": "user", "content": "Search the docs for how to send an SMS." }],
"tools": [{
"type": "function",
"function": {
"name": "search_docs",
"description": "Search the documentation.",
"parameters": {
"type": "object",
"properties": { "query": { "type": "string" } },
"required": ["query"]
}
}
}]
}'A model that can use tools answers with "finish_reason": "tool_calls" and a tool_calls entry that calls search_docs with a query. One that can't returns a 400 whose message ends in does not support tools. Passing proves the plumbing, not that the model uses tools well; the comparison later shows that.
Point the assistant at Ollama
In your Blume project, install the provider package the route uses for OpenAI-compatible endpoints:
npm install @ai-sdk/openai-compatible@3.0.57Then point the assistant at Ollama, and switch to server output with the Node adapter:
import { defineConfig } from "blume";
import { openai } from "blume/ai";
import { node } from "blume/deploy";
export default defineConfig({
ai: {
assistant: {
enabled: true,
provider: openai({
apiKeyEnv: "DOCS_LLM_API_KEY",
baseUrl: "http://127.0.0.1:11434/v1",
model: "qwen3:8b",
reasoning: "none",
}),
retrieval: {
contextBudget: 3000,
excerptChars: 1200,
maxResults: 3,
},
tools: true,
},
},
deployment: node({ site: "https://docs.acme.example" }),
});baseUrlmakesopenai()call the endpoint's Chat Completions API through@ai-sdk/openai-compatible. Without that package,blume buildstops withBLUME_DEPENDENCY_MISSING.apiKeyEnvnames the variable holding the key. Ollama ignores the key, but the route answers 503 until the variable is set.reasoning: "none"is sent asreasoning_effort, which Ollama turns into thinking off. Qwen3 thinks before every answer by default.tools: trueturns the docs tools on, because the test passed. Leave it out for a model that failed.retrievalcuts the docs sent with each question from the default 10,000 characters to 3,000.
Give the key variable any value for local development:
DOCS_LLM_API_KEY=ollamaRun npx blume dev, open the site, and ask a question in the assistant panel. The answer should stream in with links to the pages it drew from.
Run it in production next to Ollama
Build the site and start the Node server on the machine that runs Ollama, with the project's installed node_modules, since the server resolves its packages through them. .env.local is only for local development, so pass the key in the server's environment:
npx blume build
DOCS_LLM_API_KEY=ollama PORT=8080 node dist/server/entry.mjsThe server listens on localhost:8080. Run it as a service behind a reverse proxy for HTTPS; the Docker guide covers one with Caddy. In a container, 127.0.0.1 is the container itself, so use host networking, or put Ollama on the same private Docker network and point baseUrl at it.
Every question now runs on your hardware, and the route is public by design. Blume's rate limit, on by default at 30 questions per reader every 10 minutes, counts by the address the Node server sees. Behind a reverse proxy on the same machine, that's the proxy's address on every request, so all readers would share one count. Set rateLimit: false and limit /api/ask in the proxy instead, as the rate limiting guide explains. Its bot check works either way.
When the docs run somewhere else
On Vercel, Netlify, or Cloudflare, the route has to reach Ollama over the internet. Don't open port 11434 for that. Put Ollama behind a reverse proxy that serves HTTPS and checks the key the route sends as a bearer token. With Caddy, on the Ollama machine:
llm.acme.example {
@authorized header Authorization "Bearer {env.DOCS_LLM_API_KEY}"
handle @authorized {
reverse_proxy 127.0.0.1:11434 {
header_up Host {upstream_hostport}
}
}
handle {
respond 401
}
}Start Caddy with DOCS_LLM_API_KEY set to a long random value, set the same value in your docs host's environment, and change baseUrl to https://llm.acme.example/v1. Keep the header_up Host line: Ollama, listening on 127.0.0.1, answers 403 to a request whose Host header names another domain. A private network between the two machines, like a VPN, avoids the public proxy altogether.
Tune it for your hardware
On a local model, the wait for the first word is mostly the model reading its prompt, and the prompt is mostly your docs. Three things decide how long that takes and whether it works at all.
The context window
Ollama sizes the default window by GPU memory: 4k tokens under 24 GiB of VRAM, 32k up to 48 GiB, and 256k above that. Blume's prompt can outgrow 4k:
| In one request | Up to |
|---|---|
| Excerpts, including the page the reader is on | contextBudget characters (default 10,000) |
| The conversation so far | 24,000 characters |
Each search_docs result, with tools on | 5 excerpts of 700 characters |
Each read_page result, with tools on | 20,000 characters |
When a prompt doesn't fit, Ollama doesn't fail the request. It logs truncating input prompt and drops the start, where Blume's instructions and excerpts are, so the model answers without your docs. The CONTEXT column of ollama ps shows the window the model got. A bigger one takes more memory, and the PROCESSOR column shows when the model spills from the GPU onto the slower CPU.
How much documentation each question carries
The retrieval settings trade recall for speed. Start from the config above: 3 pages, 1,200 characters from each, 3,000 in total. Raise excerptChars when answers stop short of a fact on the cited page, and maxResults when the right page isn't cited at all. The comparison below tells you which.
Thinking, loading, and concurrency
Keep reasoning: "none" unless the comparison shows answers need thinking. Keep the model loaded with OLLAMA_KEEP_ALIVE, or the first question after a quiet spell waits for it to load. And know that Ollama answers one request per model at a time by default (OLLAMA_NUM_PARALLEL=1), queuing the rest. Raising it lets readers share the model at once, but memory scales with it times the context length.
Compare it with a hosted model
Keep the hosted deployment as your baseline and ask both the same questions. The hosted guide's ask-evals.json is an array of entries with an id, a question, and optionally the page it's asked from. This script asks each question on each site, one at a time, and saves every answer with the time to its first chunk and to its end. It needs Node.js 22 or later and nothing else:
import { readFile, writeFile } from "node:fs/promises";
const cases = JSON.parse(await readFile("ask-evals.json", "utf8"));
const results = [];
for (const site of process.argv.slice(2)) {
const endpoint = new URL("api/ask", site.endsWith("/") ? site : site + "/");
for (const test of cases) {
const started = performance.now();
const response = await fetch(endpoint, {
body: JSON.stringify({
messages: [{ content: test.question, role: "user" }],
page: { path: test.page ?? "/" },
}),
headers: { "content-type": "application/json" },
method: "POST",
});
const decoder = new TextDecoder();
let answer = "";
let firstChunkMs = null;
for await (const chunk of response.body ?? []) {
firstChunkMs ??= Math.round(performance.now() - started);
answer += decoder.decode(chunk, { stream: true });
}
answer += decoder.decode();
const totalMs = Math.round(performance.now() - started);
results.push({ answer, firstChunkMs, id: test.id, site, status: response.status, totalMs });
console.log([response.status, firstChunkMs ?? "-", totalMs, test.id, site].join(" "));
}
}
await writeFile("ask-compare.json", JSON.stringify(results, null, 2));Run it twice and keep the second run, so a model load doesn't count:
node ask-compare.mjs https://docs.acme.example http://localhost:8080Each line shows the status, milliseconds to the first chunk and to the end, the question, and the site. A 429 is the rate limit, so wait and rerun. A 200 with an empty answer is a model error: check the server log.
Then read ask-compare.json. For each answer, decide whether it states the facts, cites a page that holds them, and says so when the docs don't cover the question. The hosted guide's ask-eval.mjs automates the citation and wording checks: run it once per site with DOCS_URL set.
Read the misses in pairs. If both models miss, fix the page. If only the local one misses, raise retrieval first, since the baseline sends more docs, then try a larger model. If it only trails on speed, shrink retrieval or turn the tools off.
What a test run showed
We ran the pieces of this setup on an Apple M1 Pro with 16 GB of memory: Ollama 0.34.4 with qwen3:1.7b, a smaller model from the same family (ID 8f68893c685c), and Blume's retrieval and docs tools with the AI SDK packages the route imports, over Blume's own docs. It wasn't a full deployment, and your hardware will differ, but these are the behaviors to look for:
ollama psreported a 4,096-token window by default.- A question with the default retrieval came to about 3,000 tokens. After one
read_pagecall on a long page, the next request was 7,971. At 4,096, Ollama loggedtruncating input prompt, and the answer drifted off the question and cited nothing. At 16,384 it fit. - The model called
search_docsand answered from the result. The AI SDK sendstool_choice, which Ollama's docs list as unsupported; Ollama ignored it. Once, the model wrote a tool call into its answer as text instead, which a reader would see. - With thinking left on, the first chunk came later on every question than with
reasoning: "none". - It said the docs didn't cover a question they don't, and said the same about one they do, whose fact sat outside the excerpt retrieval picked: a retrieval miss, not a model one.
Troubleshooting
The panel says the assistant is not configured
The route answered 503 because DOCS_LLM_API_KEY isn't set where the server runs. .env.local only applies to blume dev, so set it in the production server's environment and restart it.
The panel says "Sorry, something went wrong."
Ollama errors arrive after the route starts streaming, so the reader gets this notice. The server log has the cause after Assistant provider error:. Cannot connect to API means Ollama isn't running or baseUrl is wrong. model 'qwen3:8b' not found means the model isn't pulled on that machine. does not support tools means the model can't call tools, so remove tools: true.
Answers ignore the docs or stop citing pages
The prompt outgrew the context window. On Linux, journalctl -u ollama | grep "truncating input prompt" shows it. Raise OLLAMA_CONTEXT_LENGTH, lower retrieval, or turn the tools off, since read_page adds the most.
A tool call shows up in the answer as text
The model wrote the call instead of making it, which small models do. Try a larger model, or remove tools: true so it answers from the retrieved excerpts alone.
The log shows a 401 or 403 from the proxy
A 401 comes from Caddy: the token on the docs host doesn't match the one Caddy was started with. A 403 comes from Ollama rejecting the Host header, so keep header_up Host {upstream_hostport}, or set the same header in whatever proxy you use.
Next step
Put a limit on your hardware
Every question runs on your own machine. Add a bot check, and limit the route at your proxy, so scripts can't queue up work for your model.
Read the rate limiting guideA step here not working for you? Report a broken step.