Agents
Test whether your documentation can answer user questions
An evals file built from real support questions, a failing run traced to the page that should answer, the fix that turns it green, and a CI gate on docs changes.
By Hayden Bleasel11 min read

To catch documentation gaps before users do, write down the questions they ask, list the facts a correct answer must contain, and have an AI agent answer them from your docs alone. In Blume, that's blume eval: a reader agent answers each question in evals.yaml using only your documentation, a second session grades the answer against your facts, and the command exits non-zero when the docs can't answer. Run it in CI and a missing fact fails the pull request instead of becoming a support ticket.
You'll end up with an evals file built from support questions, a failing run you've read, the docs change that turns it green, and a GitHub Actions job that runs the evals on docs pull requests with a pinned agent and model. The example is the docs for Acme Messages, a fictional email and SMS API at docs.acme.example.
This tests whether the agent and model you configure can answer from your docs. It doesn't measure search rankings or predict what other models say about your product, and every run uses paid model sessions. To catch broken links, blume validate needs no model and gives the same result every time (see the link checking guide). It checks the docs against questions, not against your code: to catch pages a code change made wrong, see Keep documentation updated from merged code changes. And blume eval reads a Blume project, so a site built with another tool needs its own harness.
How the eval reads your docs
Each question runs through two sessions of an agent CLI you have installed: Codex by default, or Claude Code with --agent claude. Blume itself holds no API keys.
- The reader starts in an empty directory with its file, shell, and web tools off. Its only tools are your docs, served over a private MCP server with the same
search_docs,get_page,list_pages, andget_navigationtools a real agent uses. Whatever the docs don't say doesn't exist. - The judge gets the question, your expected facts, and the answer, with no tools. A paraphrase passes. A missing or contradicted fact fails, and so does an answer that says the docs don't cover it.
The reader's snapshot comes from your content sources, so there's no build step, nothing is deployed, and the MCP server doesn't need to be enabled.
Turn support questions into expected facts
Start from questions someone asked a person instead of the docs: support threads, GitHub issues, onboarding calls, and site searches that returned nothing. For each, write down the facts your team gave in reply. For Acme, they go in evals.yaml at the project root:
questions:
- id: send-first-email
question: How do I send my first email with the Acme API?
expected:
- send a POST request to https://api.acme.example/v1/messages
- authenticate with an API key as a bearer token
- set channel to email
routes: /quickstart
- id: rotate-api-key
question: How do I rotate an API key without downtime?
expected:
- create a new key before revoking the old one
- an account can have two active keys at the same time
- revoking the old key takes effect immediately
routes: /authentication
- id: log-retention
question: How long does Acme keep message logs?
expected:
- 30 days
routes: /limits
- id: scheduled-send
question: Can I schedule a message to send later?
expected:
- set send_at to an ISO 8601 timestamp
severity: warningexpectedlists one checkable fact per line. Write the substance, not the wording: a value (30 days), a name (send_at), an order (create before revoke). A fact like "explains authentication clearly" gives the judge nothing to check.routesnames the page that should answer, so a failure is reported against that page's source file.severity: warningreports a miss without failing the run. Scheduled sending ships soon and isn't documented yet, so it shouldn't block anyone.skip: truesits a question out entirely.
rotate-api-key is the one to watch: support answers it every week, and the authentication page never says two keys can be active at once.
Run the evals
This guide uses Claude Code because it's the agent whose model you can pin: it reads ANTHROPIC_MODEL. Blume runs Codex with --ignore-user-config and passes it no model, so Codex uses its default. Install Claude Code and sign in (or set ANTHROPIC_API_KEY), then run the evals from the project root:
npm install -g @anthropic-ai/claude-code@2.1.283
export ANTHROPIC_MODEL=claude-sonnet-5
npx blume eval --agent claude --json > eval-report.jsonProgress goes to your terminal, and --json writes the full report, answers included, to the file. Each question's line shows its status, the judge's score from 0 to 1, its duration, and the cost estimate. Missing facts are listed underneath, and a fix: line at the end names the source file to edit.
Read a failure
In the run behind this guide, two of the four questions missed and the command exited with code 1. Pull the misses out of the report:
jq '.eval.results[] | select(.status != "pass") | {id, score, missing, notes}' eval-report.json{
"id": "rotate-api-key",
"score": 0.1,
"missing": [
"create a new key before revoking the old one",
"an account can have two active keys at the same time"
],
"notes": "The answer explicitly claims the documentation lacks this information and fails to state any of the expected facts, instead framing them as unknown/missing."
}
{
"id": "scheduled-send",
"score": 0,
"missing": [
"set send_at to an ISO 8601 timestamp"
],
"notes": "The answer claims scheduling is undocumented and unsupported, directly contradicting the expected fact that a send_at ISO 8601 parameter exists."
}rotate-api-key failed, reported against docs/authentication.mdx, the file behind its route. The judge credited the third fact, because the page does say revoking takes effect immediately:
---
title: Authentication
description: Authenticate Acme Messages API requests with an API key.
---
Every request needs an API key, sent as a bearer token:
```http
Authorization: Bearer YOUR_API_KEY
```
Create keys in the dashboard under **Settings > API keys**. A key is shown
once, when you create it, so store it in your secrets manager right away.
## Revoke a key
Select **Revoke** next to the key. Revoking takes effect immediately, and
requests that use the key get a `401` response.The reader's full answer is in the report's answer field, or under each failure with --verbose. Here it said the documentation "does not describe a zero-downtime rotation process" and has "no mention of support for multiple simultaneous active keys." It even sketched the usual rotation steps from general knowledge, and still failed: the grade is about what your docs support, not what a model knows.
scheduled-send missed too, but as a warning question it doesn't fail the run. With no route, its finding points at its line in evals.yaml.
Fix the docs, not the question
Add the missing facts to the page the finding names, in its own voice. For Acme, that's a new section above "Revoke a key":
## Rotate a key
An account can have two active keys at the same time, so you can rotate a key
without downtime:
1. Create a new key under **Settings > API keys**.
2. Deploy the new key everywhere the old one is used.
3. Revoke the old key.Run the evals again. All three error-level questions pass and the command exits 0. scheduled-send still misses: the terminal marks it failed and counts it in the "failed" total, but it doesn't touch the exit code, and the JSON report's summary counts it as a warning.
Don't edit the question to get to green. If a fact is wrong because the product changed, update the fact and the docs in the same pull request. To hand the fixes to an agent, run npx blume eval --fix --agent claude. After the run, it opens Claude Code with the report and instructions to add the missing facts to the named pages, never deleting a question or weakening a fact. Those instructions end by rerunning plain blume eval, which uses Codex, so tell the agent to add --agent claude, or rerun the evals yourself with your CI flags once you've reviewed its edits.
Account for model variability
Two model sessions stand between your docs and each verdict, and neither is deterministic: the reader can search differently from run to run, and the judge can weigh a paraphrase differently. A gate is useful only when it fails because the docs changed, so control what you can.
Pin what the result depends on
- Blume, through your lockfile.
- The agent CLI, installed at an exact version.
- The model, as a full model ID in
ANTHROPIC_MODEL. An alias likesonnetmoves to a newer model when one ships. - The evals file, committed beside the docs, so changes to questions and facts get reviewed.
Blume doesn't start Claude Code in bare mode, so a local run also loads your own ~/.claude: a model in your user settings (which ANTHROPIC_MODEL overrides) and a personal CLAUDE.md. CI runners have none of that.
Check that a new question is stable
Before a question joins the gate, run the suite a few times against the same docs and count the outcomes:
export ANTHROPIC_MODEL=claude-sonnet-5
for i in 1 2 3; do
npx blume eval --agent claude --json > "eval-$i.json"
done
jq -r '.eval.results[] | "\(.id) \(.status)"' eval-*.json | sort | uniq -cEach line counts the runs that gave a question a status. A question that both passes and fails on unchanged docs is noise: usually the page states the fact only indirectly, or the fact itself is vague. Make whichever it is explicit. Keep the question at severity: warning until it's steady.
Know what a run costs
Every question that isn't skipped is two agent sessions, run one after another, so cost and time grow with the file. The sessions bill to whatever the agent CLI signs in with: your subscription or key locally, the API key secret in CI. With Claude Code, Blume prints its estimate after each question and a total on the summary line, and stores the total as eval.costUsd in the JSON report. Claude Code calls these client-side estimates that can differ from your bill. Codex reports none, so check your provider's usage dashboard.
Keep the file to questions that matter, and run it only when the docs change. The pricing page covers what Blume's AI features cost.
Add a CI gate on docs changes
Add your Anthropic API key as a repository secret named ANTHROPIC_API_KEY, then add this workflow. It assumes the Blume project is at the repository root with its content in docs/ and a package-lock.json; adjust the paths and install command to match yours.
name: Docs evals
on:
pull_request:
paths:
- "docs/**"
- "evals.yaml"
- "blume.config.ts"
workflow_dispatch:
permissions:
contents: read
concurrency:
group: docs-evals-${{ github.ref }}
cancel-in-progress: true
jobs:
evals:
# Pull requests from forks get no secrets, so the job skips them.
if: github.event_name == 'workflow_dispatch' || github.event.pull_request.head.repo.full_name == github.repository
runs-on: ubuntu-latest
timeout-minutes: 30
steps:
- uses: actions/checkout@v7
- uses: actions/setup-node@v7
with:
node-version: 24
- run: npm ci
- run: npm install -g @anthropic-ai/claude-code@2.1.283
- name: Run the evals
env:
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
ANTHROPIC_MODEL: claude-sonnet-5
run: npx blume eval --agent claude --json > eval-report.json
- name: Summarize
if: always()
run: |
[ -s eval-report.json ] || exit 0
{
echo "## Docs evals"
jq -r '.eval.results[] | "- \(.id): \(.status)"' eval-report.json
jq -r '"Cost estimate (USD): \((.eval.costUsd // 0) * 100 | round / 100)"' eval-report.json
} >> "$GITHUB_STEP_SUMMARY"
- uses: actions/upload-artifact@v7
if: always()
with:
name: eval-report
path: eval-report.jsonEach part keeps the gate scoped:
- The
pathsfilter runs the evals only when the docs, the questions, or the config change. - The fork check skips pull requests from forks, which GitHub gives no secrets. A skipped job reports success, so run the evals yourself before merging a fork's docs change.
- The API key is set on the eval step alone, so install scripts that run during
npm cinever see it. concurrencycancels a run when a newer push supersedes it.- The report is uploaded on every run, and the job summary lists each question's status and the cost estimate.
timeout-minutescaps the job. The reader gets 180 seconds per question by default (--timeout), so raise the cap as the file grows.
Leave the check optional at first. GitHub leaves a required check pending when a paths filter skips its workflow, so if you make it required, pull requests that don't touch the docs can't merge. Keep it as a signal reviewers read, or drop the filter and run it on every pull request.
When you add evals to docs with many known gaps, mark those questions severity: warning rather than lowering the bar with --threshold 0.8. A threshold lets any question fail as long as enough others pass; warnings record which gaps you already know about.
Troubleshooting
The agent CLI isn't found
Blume stops with "Claude Code (claude) was not found on PATH" and the install command. In CI, the install step has to come before the eval step. Without --agent claude, Blume looks for Codex instead.
A question says "run failed" instead of fail
BLUME_EVAL_QUESTION_ERROR means the agent session itself failed, so the docs weren't graded: the reader timed out, the CLI exited with an error, or the judge returned no parseable verdict. The detail line says which. Check the API key for authentication errors, raise --timeout for slow readers, and rerun. An errored question counts against the gate unless it's a warning question.
The fact is on the page, but the question still fails
The reader finds pages through search_docs, so check that it can. A page with search.exclude in its frontmatter, or one hidden from the sidebar, is left out of that index (see Excluding pages). Code blocks aren't indexed unless you turn on search.indexing.includeCodeBlocks, so the reader can miss a fact that only appears in an example; state it in prose too. Content inside <Visibility for="web"> never reaches agents. If users ask in words the page doesn't use, add those words to the page.
A route hint matches no page
BLUME_EVAL_ROUTE_UNKNOWN warns that a question's routes entry names a page that moved or no longer exists. Update it to the new route; until then, failures point at the question in evals.yaml instead of the page.
The evals file won't load
"Invalid evals file" names the field at fault: an id that isn't kebab-case, a duplicate id, an empty expected list, or a key the schema doesn't know. "No evals file found" means Blume looked for evals.yaml in the current directory; run it from the project root or pass --file.
It passes locally and fails in CI
Check that both runs used the same ANTHROPIC_MODEL and CLI version, and that your own Claude Code settings aren't the difference. If they match, the question may be unstable: run the stability check above and compare the reader's answer in the CI artifact with your local report.
Next step
Draft your first evals file
Have an agent draft questions from your docs, then add the ones your users actually asked.
npx blume eval init --agent claudeA step here not working for you? Report a broken step.