Writing
Build an engineering handbook with RFCs, ADRs, and runbooks
A handbook of RFCs, decision records, and runbooks where the build rejects a missing owner or status, and agents can list documents by type, team, or status.
By Hayden Bleasel12 min read

Keep your RFCs, architecture decision records (ADRs), and runbooks as Markdown files in one Git repository, and give each kind its own frontmatter type with the fields it must carry, like an owner and a status. Blume builds them into one searchable site, fails the build when a document is missing its metadata, and serves an MCP endpoint where an agent can ask for only the accepted decisions a team owns.
By the end, you have the Acme Engineering Handbook: three sections with a card index each, validated owner, status, and service fields, and an MCP server you can query by content type and by those fields. The example holds one RFC, two decisions, and one runbook, so you can see every piece working before you move your own documents in.
Blume publishes and indexes the handbook. It doesn't run your process: pull requests stay the place where RFCs get reviewed and approved. It also has no sign-in or per-page permissions. The whole site is private or public at your host, so if some pages must be hidden from some readers, publish them as a separate site or use a tool with page-level permissions.
Plan the content types
Every Blume page has a type in its frontmatter, and doc is the default. You can invent your own types and attach required fields to each one. Fields you list as facets become filters an agent can use. The handbook uses three:
| Type | Folder | Required fields | Facets |
|---|---|---|---|
rfc | rfcs/ | owner, status | owner, status |
adr | decisions/ | owner, status | owner, status |
runbook | runbooks/ | owner, service, reviewed | owner, service |
The folder decides where a page appears in the sidebar, and the type decides which fields it must have. They're independent, so every page states its type even though the folder already hints at it. RFCs and decisions both use status, with different values, and both include accepted, so one filter can find accepted documents of either kind.
Create the project
Scaffold a Blume project and add Zod for the field schemas:
npx blume init acme-handbook --template docs
cd acme-handbook
npm install zod@4.6.5Keep the default docs content folder when init asks. Any Standard Schema library works for the schemas (Valibot and ArkType too); this guide uses Zod. By the end, the project looks like this:
acme-handbook/
├── blume.config.ts
├── package.json
└── docs/
├── index.mdx
├── rfcs/
│ ├── meta.ts
│ ├── index.mdx
│ └── 0001-idempotency-keys-for-message-sends.mdx
├── decisions/
│ ├── meta.ts
│ ├── index.mdx
│ ├── 0001-retry-webhooks-at-a-fixed-interval.mdx
│ └── 0002-retry-webhooks-with-exponential-backoff.mdx
└── runbooks/
├── meta.ts
├── index.mdx
└── webhook-delivery-backlog.mdxDefine the types and their required fields
Replace blume.config.ts with:
import { defineConfig } from "blume";
import { z } from "zod";
// The teams that can own a document. A typo on any page fails the build.
const owner = z.enum(["platform", "payments", "messaging"]);
export default defineConfig({
title: "Acme Engineering Handbook",
description:
"How Acme's engineers propose changes, record decisions, and run production.",
content: {
types: {
rfc: {
facets: ["owner", "status"],
frontmatter: {
owner,
status: z.enum(["draft", "review", "accepted", "rejected", "withdrawn"]),
},
},
adr: {
facets: ["owner", "status"],
frontmatter: {
owner,
status: z.enum(["proposed", "accepted", "deprecated", "superseded"]),
},
},
runbook: {
facets: ["owner", "service"],
frontmatter: {
owner,
service: z.string().regex(/^[a-z0-9-]+$/),
reviewed: z.coerce.date(),
},
},
},
},
});Here's what each part enforces:
- Every key under a type's
frontmatteris checked on every page of that type, including when it's missing. A decision without anownerfails, and so doesstatus: acepted. See Per-type keys for the full rules. - The keys stay unknown on every other type, so a stray
statuson the home page also fails. Blume's frontmatter is strict everywhere, and these keys are the only additions. facetsmust name keys the type declares. Only string, number, and boolean values become facets, soreviewed, a date, is checked but can't be filtered on.- Facet filters match exact strings. Enums and a lowercase pattern for
servicekeep values likeWebhooksandwebhooksfrom splitting one service in two.
An RFC's status: draft is your own field and has nothing to do with Blume's built-in draft: true, which leaves a page out of production builds entirely. An RFC in draft should still be published, so people can comment on it.
Write the documents
An RFC
---
title: "RFC 1: Idempotency keys for message sends"
description: Let clients retry POST /messages without sending the same email or SMS twice.
type: rfc
owner: messaging
status: accepted
---
## Problem
When a client's request to `POST /messages` times out, the client can't tell whether we sent the message. Retrying risks a duplicate email or SMS.
## Proposal
Accept an `Idempotency-Key` header on `POST /messages`. Store each key with the response for 24 hours, and return the stored response when the same key arrives again.
## Alternatives considered
Deduplicating on the message body would block legitimate repeat sends, like two identical password reset codes.
## Open questions
None. Accepted in the platform review on 2026-08-12.Two decisions, one superseding the other
Decision records are never edited after they're accepted. When a decision changes, a new record replaces it and the old one is marked superseded. Both records link to each other through related, which shows them as cards at the foot of each page:
---
title: "ADR 1: Retry webhooks at a fixed interval"
description: Retry failed webhook deliveries every 60 seconds, up to 10 times.
type: adr
owner: messaging
status: superseded
related:
- /decisions/retry-webhooks-with-exponential-backoff
---
<Callout type="info" title={`Status: ${frontmatter.status}`}>
A record doesn't change once it's accepted. To reverse it, write a new record
that supersedes it.
</Callout>
## Context
Customer endpoints fail for short periods, and we need a retry policy before launch.
## Decision
Retry each failed delivery every 60 seconds, up to 10 times.
## Consequences
Simple to build. Superseded by ADR 2 after a customer outage turned our retries into a flood.---
title: "ADR 2: Retry webhooks with exponential backoff"
description: Back off exponentially between webhook retries, with jitter, for up to 24 hours.
type: adr
owner: messaging
status: accepted
related:
- /decisions/retry-webhooks-at-a-fixed-interval
---
<Callout type="info" title={`Status: ${frontmatter.status}`}>
A record doesn't change once it's accepted. To reverse it, write a new record
that supersedes it.
</Callout>
## Context
When a customer's endpoint was down for an hour, ADR 1's fixed 60-second retries sent a burst of deliveries the moment it recovered.
## Decision
Double the delay after each failed attempt, starting at 30 seconds, with up to 20% random jitter. Stop after 24 hours and mark the delivery failed.
## Consequences
Recovering endpoints see a gradual ramp instead of a burst. A delivery can now arrive hours late, so the webhook docs say so.Blume validates your custom fields but doesn't print them on the page. The callout reads the status from the frontmatter, so it can't drift from the field agents filter on, and the page's Markdown copy for agents shows the same "Status: accepted". Keep frontmatter references in component props like title: an expression in a component's body isn't resolved in that Markdown copy.
The 0001- prefix sorts the records in the sidebar and is dropped from the URL, so ADR 2 lives at /decisions/retry-webhooks-with-exponential-backoff. Keep the number in the title, where people cite it.
A runbook
---
title: Webhook delivery backlog
description: The webhook queue is growing faster than workers drain it.
type: runbook
owner: messaging
service: webhooks
reviewed: 2026-09-01
---
## Alert
`WebhookQueueDepthHigh` fires when more than 50,000 deliveries wait for over 10 minutes.
## Check
1. Open the webhooks dashboard and find the endpoints with the most pending deliveries.
2. If one endpoint holds most of the backlog, it's that customer's outage, not ours. Backoff handles it (see ADR 2).
3. If the backlog is spread across endpoints, check the worker error rate.
## Fix
- Scale the webhook workers: `kubectl scale deployment webhook-worker --replicas=12 -n messaging`.
- If workers are crashing, roll back the last webhooks deploy.
## Escalate
Page the messaging team's secondary on-call if the queue is still growing 30 minutes after scaling.When several runbooks repeat the same steps, like how to page the secondary on-call, keep those steps in one partial and include it in each runbook, as the snippets guide shows.
Build the navigation
Blume builds the sidebar from your folders, one group per folder. Give each group a label and an icon, and turn its landing page into an index of the documents inside it, with a meta.ts:
import { defineMeta } from "blume";
export default defineMeta({
title: "Decisions",
icon: "scale",
directory: "card",
});Add the same file to the other two folders, with title: "RFCs" and icon: "message-square-text" in docs/rfcs/meta.ts, and title: "Runbooks" and icon: "siren" in docs/runbooks/meta.ts. Without the title, the RFCs group would read "Rfcs".
directory: "card" lists the group's pages as cards, each with its title and description, below the folder's index page (see Directory listings). The listing needs that page to show on, so give each folder one:
---
title: Decisions
description: Architecture decision records. One decision per page, never edited after it's accepted.
---
Each record captures one decision, the context it was made in, and its consequences. To change a decision, write a new record that supersedes it, and set the old one's `status` to `superseded`.---
title: RFCs
description: Proposals for changes that cross team boundaries, open for comment before anyone builds them.
---
Write an RFC before a change that affects another team's service, a public API, or how we store customer data. Copy an existing RFC, set `status: draft`, and open a pull request.---
title: Runbooks
description: What the on-call engineer does when an alert fires, one runbook per alert.
---
Each runbook matches one alert. If an alert has no runbook, write one after the incident.Finally, replace the home page:
---
title: Acme Engineering Handbook
description: How Acme's engineers propose changes, record decisions, and run production.
---
This handbook holds three kinds of document:
- [RFCs](/rfcs) propose a change and collect feedback before anyone builds it.
- [Decisions](/decisions) record what we chose and why, one decision per page.
- [Runbooks](/runbooks) tell the on-call engineer what to do when something breaks.
Every RFC, decision, and runbook names the team that owns it.Run npx blume dev and open http://localhost:4321. The sidebar has Decisions, RFCs, and Runbooks groups, and each group's page ends in a card per document. For a long decision log, try directory: "accordion", which lists pages as rows instead. Search covers every page, whatever its type.
Query by type and facet through MCP
Humans browse the sidebar. Agents do better with a filter: "list the accepted decisions the messaging team owns" should be one call, not a crawl. Turn on Blume's MCP server, and name a host adapter, because the server is a live endpoint and a static build can't include it. Add the import and the two new keys to blume.config.ts:
import { defineConfig } from "blume";
import { vercel } from "blume/deploy"; // or netlify, cloudflare, node
import { z } from "zod";
// ...the owner schema from above
export default defineConfig({
// ...title, description, and content from above
agents: {
mcp: {
enabled: true,
instructions:
"Pages have a content type: rfc (proposals), adr (architecture decisions), or runbook (on-call steps). Scope search_docs and list_pages with contentTypes, and with filters on owner, status (rfc and adr), or service (runbook). Facet values are lowercase.",
},
},
deployment: vercel(),
});instructions is passed to every agent that connects, so it learns your types and facets before its first call. The MCP server guide covers the adapter choice and deployment in detail.
Try a filter
With npx blume dev running, the server answers at http://localhost:4321/mcp. Ask list_pages for accepted decisions. The last line uses jq to unwrap the tool's text result; drop it to see the raw JSON-RPC response.
curl -s http://localhost:4321/mcp \
-H "Content-Type: application/json" \
-H "Accept: application/json, text/event-stream" \
-d '{"jsonrpc":"2.0","id":1,"method":"tools/call","params":{"name":"list_pages","arguments":{"contentTypes":["adr"],"filters":{"status":"accepted"}}}}' \
| jq -r '.result.content[0].text'Only ADR 2 comes back, with its facet values:
[
{
"contentType": "adr",
"description": "Back off exponentially between webhook retries, with jitter, for up to 24 hours.",
"facets": {
"owner": "messaging",
"status": "accepted"
},
"lastModified": null,
"route": "/decisions/retry-webhooks-with-exponential-backoff",
"title": "ADR 2: Retry webhooks with exponential backoff",
"url": "http://localhost:4321/decisions/retry-webhooks-with-exponential-backoff"
}
]Swap the arguments object to ask other questions. search_docs takes the same contentTypes and filters beside its query:
| Question | Tool | Arguments |
|---|---|---|
| Every decision record, with its status | list_pages | {"contentTypes": ["adr"]} |
| Accepted RFCs and decisions | list_pages | {"filters": {"status": "accepted"}} |
| Runbooks for the webhooks service that mention a backlog | search_docs | {"query": "queue backlog", "contentTypes": ["runbook"], "filters": {"service": "webhooks"}} |
Every entry in filters must match, and a page whose type doesn't declare a key never matches a filter on it. That's why the second row returns RFC 1 and ADR 2 but never a runbook. To read a document in full, an agent passes its route to get_page, which returns the Markdown with its frontmatter, so the owner and status come along. The MCP docs list every argument the tools take.
The same page list, facets included, is also a static file at /api/docs/pages.json, part of the JSON API, for scripts that don't speak MCP.
Connect Claude Code
From the handbook repository, register the local server for everyone who clones it:
claude mcp add --transport http --scope project acme-handbook http://localhost:4321/mcp--scope project writes the server to a .mcp.json you commit. Anyone with the repository runs npx blume dev, and their agent can ask "Which accepted decisions does the messaging team own?" and answer it with one list_pages call.
Keep the handbook private
An engineering handbook is usually internal, and Blume has no sign-in of its own. Privacy comes from your host, which checks every request before it serves anything. On Vercel, turn on Deployment Protection with the All Deployments scope, so production is covered and not only previews; Vercel Authentication then lets in members of your Vercel team. Private docs lists the equivalent on Netlify, Cloudflare, GitHub Pages, and your own server. Protection covers the pages, the search index, the Markdown copies, llms.txt, and the MCP endpoint alike.
Host protection also stops agents: a client that can't sign in gets a login page, not the MCP server. The local server from the last section avoids that, since the repository's permissions decide who can run it. To serve agents from the deployed site instead, you need a host that accepts credentials from a non-browser client. Cloudflare Access does, with a service token and a policy whose action is Service Auth. Claude Code can send the token's two headers:
claude mcp add --transport http acme-handbook https://handbook.acme.example/mcp \
--header "CF-Access-Client-Id: $CF_ACCESS_CLIENT_ID" \
--header "CF-Access-Client-Secret: $CF_ACCESS_CLIENT_SECRET"Leave this one at the default local scope, so the token stays in your own Claude Code settings and out of the repository. The Cloudflare Access guide walks through protecting a Blume site on Cloudflare. And if you switch to a hosted search provider later, remember that it keeps its own copy of your content outside the protection. The built-in search index stays in your build.
Troubleshooting
The build fails with BLUME_FRONTMATTER_INVALID
A page is missing a required field or has a value the schema rejects. The diagnostic names the file and the field, such as owner: Invalid option: expected one of "platform"|"payments"|"messaging". Fix the page, or add the new value to the enum if it's a new team. blume dev reports the same error and leaves the page out of the site, so a page that's missing in dev is worth checking in the terminal.
A page fails with "Unrecognized keys"
The page's type is misspelled, like type: runbok. Blume then treats it as a type with no fields of its own, so owner, service, and reviewed are unknown keys. Correct the type.
The config fails with BLUME_CONFIG_INVALID on a facet
A name in facets isn't a field the type declares, as in Facet "severity" for type "runbook" is not a declared custom frontmatter key. Add the field to the type's frontmatter map, or remove it from facets.
A filter returns nothing, or everything
Nothing: facet values match exactly and are case-sensitive, so "Accepted" doesn't find accepted. Check the values list_pages reports. Everything: only string values count as filters, and any other value is dropped without an error, so {"status": 1} filters nothing. Send numbers and booleans as strings, like "1" or "true".
A document doesn't show up in list_pages or search_docs
First check the terminal for a frontmatter error, which leaves the page out of the site entirely. Then check its frontmatter: sidebar.hidden: true drops it from list_pages, and from search_docs unless search.indexing.includeHiddenPages is on, while search.exclude: true drops it from search_docs only. Either way, get_page still returns it.
The build fails with BLUME_SERVER_FEATURE_REQUIRED
The MCP server is on, but the build is static. Set deployment to a host adapter from blume/deploy, and if it has output: "static", remove that.
Next step
Start your handbook
Scaffold a project, add the three content types to its config, and move your first decision record in.
npx blume init acme-handbook --template docsA step here not working for you? Report a broken step.