Skip to content
Blume
Esc
↑↓navigate↵open⌘Jpreview
Guides

Agents

Serve HTML and Markdown from the same documentation URL

Browsers get HTML and agents that send Accept: text/markdown get the page's Markdown at the same URL, checked by a script that tests headers, content, and caching.

By 12 min read

Yes. A Blume site already serves every content page twice: as HTML at its URL, and as Markdown at the same URL with .md on the end. Deploy it as a server build on Vercel or Cloudflare, and the page URL itself answers with Markdown when a request's Accept header asks for text/markdown. Browsers keep getting HTML, because they never ask for Markdown.

By the end of this guide, your docs negotiate at every content URL, and you have a script that proves it: it checks both formats' headers and bodies and the page's canonical URL, and catches a cache that hands one client the other's format.

You may not need negotiation at all. Every Blume build, static ones included, publishes the .md URLs and points to each one from the page's HTML, so an agent that looks for them finds them on any host. Negotiation helps agents that request the page URL they were given and say which format they want, and it needs a host where Blume can act on each request. To pull every page at once, for a search index or a RAG pipeline, use the JSON API instead, as in Index documentation for RAG without scraping HTML.

How one URL serves two formats

Every content page has three addresses:

URLReturns
/send-a-messageThe rendered page, or its Markdown when the request asks for it
/send-a-message.mdMarkdown, with components converted to plain Markdown
/send-a-message.mdxThe MDX source, with components as written

The homepage's Markdown lives at /index.md. When the homepage is a custom landing page rather than a content page, its Markdown is the site's llms.txt index, so an agent asking the site root for Markdown still gets a map of the docs.

A request gets Markdown when its Accept header lists text/markdown (or text/x-markdown) at least as high as text/html (Vercel skips that comparison: see Troubleshooting). A browser opening a page sends text/html first and never mentions Markdown. The answer is a rewrite, not a redirect: the agent stays on the page URL, and the body is the same file the .md URL serves.

Where that decision runs depends on how you deploy:

Where the site runsMarkdown at the page URLMarkdown at .md URLs
blume devYesYes
vercel() server buildYes, through routing rules Blume adds to the deployYes
cloudflare() server buildYes, through a Worker Blume generates in front of the siteYes
Static builds on any host, netlify(), and node()NoYes

The last row isn't a missing setting. Those hosts serve prerendered pages from a static layer that runs no code per request, so there's nowhere for the decision to happen, and agents there use the .md URLs. An adapter with output: "static" is a static build too. If you're still choosing, see static or server-rendered documentation.

Look at the Markdown mirror

Start with the dev server, which negotiates the same way. Take a page like this one, saved as docs/send-a-message.mdx:

---
title: Send your first message
description: Send a transactional email with the Acme Messages API and check that it arrived.
---

Every message goes through `POST /messages`, authenticated with an API key.
Create one first, as [Authentication](/authentication) describes.

:::warning
Keep API keys on your server. Never ship one in a browser bundle.
:::

## Send the message

<Steps>
  <Step title="Create a template">
    Add a template named `welcome` in the dashboard.
  </Step>
  <Step title="Call the API">
    Send the recipient and the template name.
  </Step>
</Steps>

```bash
curl https://api.acme.example/v1/messages \
  -H "Authorization: Bearer $ACME_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"channel":"email","to":"ada@example.com","template":"welcome"}'
```

## Check delivery

Fetch the message by its ID until its `status` is `delivered`.

Run npx blume dev, then fetch the page's Markdown both ways from another terminal:

curl -s http://localhost:4321/send-a-message.md
curl -s -H "Accept: text/markdown" http://localhost:4321/send-a-message

Both print the same Markdown. The front matter stays, so an agent gets the title and description, and plain Markdown passes through as written, the :::warning directive included. Only components change. Here, <Steps> becomes a numbered list:

## Send the message

1. **Create a template**

    Add a template named `welcome` in the dashboard.

2. **Call the API**

    Send the recipient and the template name.

The HTML side of the pair is the same text inside your site's layout, with the header, sidebar, and styled components around it. Its head ties the two together. With the site URL set to https://docs.acme.example, the page carries:

<link rel="canonical" href="https://docs.acme.example/send-a-message">
<link href="/send-a-message.md" rel="alternate" type="text/markdown">

The canonical link names the page URL as the one identity for both formats. The alternate link tells any client where the Markdown lives, even on a host that can't negotiate. Your sitemap lists page URLs only, never the .md files.

Deploy to a host that negotiates

Name your host's adapter as deployment in blume.config.ts. That switches the build to server output, which is what turns negotiation on. Your content pages are still prerendered, and the host decides per request which prerendered file to serve.

Vercel

import { defineConfig } from "blume";
import { vercel } from "blume/deploy";

export default defineConfig({
  deployment: vercel(),
});

The vercel() adapter ships with Blume. Push the repository, import it in Vercel, and set the project's Node.js version to 22 or later. Blume detects your production domain for the canonical URLs, and adds header-conditional rewrites to the deploy's routing config, so Vercel's router sends a Markdown request to the page's .md file.

Cloudflare Workers

Install the adapter, which Blume doesn't bundle:

npm install -D @astrojs/cloudflare@14

Then name it with your site URL. Blume can't detect your domain on a Workers build, and without a site URL your pages have no canonical link.

import { defineConfig } from "blume";
import { cloudflare } from "blume/deploy";

export default defineConfig({
  deployment: cloudflare({ site: "https://docs.acme.example" }),
});

Build and deploy from the project root. Run npx wrangler login first if you haven't.

npx blume build
npx wrangler deploy

Blume generates a small Worker in front of the site's own. It serves the page's .md file when a request prefers Markdown and passes everything else through. For that to work, every page request goes through the Worker, while the build assets and the raw .md, .txt, and .well-known files stay on Cloudflare's static layer. The Cloudflare Workers guide covers custom domains and the rest of the setup.

Confirm the build wired it

The build log says when negotiation is in place, with one of these lines:

Wired Accept: text/markdown negotiation into the Vercel routing config
Wired Accept: text/markdown negotiation into the Cloudflare Worker

Once deployed, the site's agent readability manifest says so too. Its artifacts.markdown entry lists contentNegotiation only on a deployment that honors the header:

curl -s https://docs.acme.example/agent-readability.json
"markdown": {
  "contentNegotiation": "text/markdown",
  "pattern": "https://docs.acme.example/{route}.md"
}

Send both Accept headers

Ask for the same URL twice, once as a browser and once as an agent, and print the two headers that matter:

curl -sI https://docs.acme.example/send-a-message \
  | grep -iE '^(content-type|vary):'
curl -sI -H "Accept: text/markdown" https://docs.acme.example/send-a-message \
  | grep -iE '^(content-type|vary):'

The first should report text/html and the second text/markdown, and both should list Accept in Vary. That header tells caches the URL has two representations. Your host may add other values to it, like accept-encoding. On Cloudflare, the Worker sends the Markdown as text/markdown; charset=utf-8.

On both hosts, missing pages negotiate too. A request for a URL with no page behind it that asks for Markdown gets a short Markdown not-found page, with links to each top-level section, the sitemap, and llms.txt, and a real 404 status. A 404 page of your own replaces it, so missing pages then answer with your HTML page.

Compare the two representations

Headers only prove the format. To check that both formats describe the same page, save this script as check-negotiation.mjs. It needs Node.js 22 or later and no dependencies.

// Checks that each docs URL serves HTML to a browser and Markdown to an agent.
// Usage: node check-negotiation.mjs <site> <path> [path...]
const [site, ...paths] = process.argv.slice(2);
if (!site || paths.length === 0) {
  console.error("Usage: node check-negotiation.mjs <site> <path> [path...]");
  process.exit(2);
}

// What Firefox sends when you open a page, and what an agent sends.
const BROWSER =
  "text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8";
const AGENT = "text/markdown";
const CACHE_HEADERS = ["x-vercel-cache", "cf-cache-status", "x-cache", "age"];

let failures = 0;
const report = (ok, message) => {
  if (!ok) failures += 1;
  console.log(`  ${ok ? "ok  " : "FAIL"} ${message}`);
};

const get = async (url, accept) => {
  const response = await fetch(url, { headers: { accept } });
  return {
    body: await response.text(),
    headers: response.headers,
    status: response.status,
    type: response.headers.get("content-type") ?? "(none)",
  };
};

const variesOnAccept = (headers) =>
  (headers.get("vary") ?? "")
    .toLowerCase()
    .split(",")
    .some((value) => value.trim() === "accept");

const cacheInfo = (headers) =>
  CACHE_HEADERS.filter((name) => headers.has(name))
    .map((name) => `${name}: ${headers.get(name)}`)
    .join(", ");

const decode = (text) =>
  text
    .replaceAll("&lt;", "<")
    .replaceAll("&gt;", ">")
    .replaceAll("&quot;", '"')
    .replaceAll("&#39;", "'")
    .replaceAll("&#x27;", "'")
    .replaceAll("&amp;", "&");

// Blume typesets quotes and dashes in the HTML, so compare plain punctuation.
const squash = (text) =>
  text
    .replace(/[\u2018\u2019]/gu, "'")
    .replace(/[\u201C\u201D]/gu, '"')
    .replace(/\u2026/gu, "...")
    .replace(/-{2,3}|[\u2013\u2014]/gu, "-")
    .replace(/\s+/gu, " ")
    .trim();

// Markdown lines outside fenced code blocks.
const prose = (markdown) => {
  let fenced = false;
  return markdown.split("\n").filter((line) => {
    if (/^\s*(```|~~~)/u.test(line)) {
      fenced = !fenced;
      return false;
    }
    return !fenced;
  });
};

// Headings, minus the anchor and TOC markers (`[#id]`, `[!toc]`); a `[toc]`
// heading renders only in the table of contents, so it's skipped.
const markdownHeadings = (markdown) =>
  prose(markdown)
    .map((line) => /^#{2,6}\s+(.+?)\s*#*$/u.exec(line)?.[1])
    .filter((text) => text && !/\[toc\]/u.test(text))
    .map((text) => text.replace(/\s*(\[(#[^\]]+|!toc)\]|\\?\{#[^}]*\})/gu, ""))
    .map((text) =>
      squash(
        text
          .replace(/\[([^\]]*)\]\([^)]*\)/gu, "$1")
          .replaceAll("`", "")
          .replaceAll("**", ""),
      ),
    );

const htmlHeadings = (html) =>
  [...html.matchAll(/<h[1-6]\b[^>]*>([\s\S]*?)<\/h[1-6]>/giu)].map(
    ([, inner]) => squash(decode(inner.replace(/<[^>]+>/gu, ""))),
  );

// Root-relative page links in the prose: not images, not inline code.
const markdownLinks = (markdown) => {
  const links = new Set();
  for (const line of prose(markdown)) {
    const text = line.replace(/`[^`]*`/gu, "");
    for (const [, target] of text.matchAll(/(?<!!)\[[^\]]*\]\((\/[^)\s]*)/gu)) {
      links.add(target);
    }
  }
  return [...links];
};

const htmlLinks = (html) =>
  new Set(
    [...html.matchAll(/\shref="([^"]*)"/giu)].map(([, href]) => decode(href)),
  );

const linkTags = (html) =>
  [...html.matchAll(/<link\b[^>]*>/giu)].map(([tag]) =>
    Object.fromEntries(
      [...tag.matchAll(/([a-z-]+)="([^"]*)"/giu)].map(([, name, value]) => [
        name.toLowerCase(),
        decode(value),
      ]),
    ),
  );

for (const path of paths) {
  const url = new URL(path, site).href;
  const mirror = new URL(
    path === "/" ? "/index.md" : `${path.replace(/\/$/u, "")}.md`,
    site,
  ).href;
  console.log(path);

  // Browser, then agent, then browser again, so a cache that ignores Accept
  // shows up as the wrong format on one of them.
  const html = await get(url, BROWSER);
  const markdown = await get(url, AGENT);
  const again = await get(url, BROWSER);
  const file = await get(mirror, "*/*");

  report(
    html.status === 200 && html.type.startsWith("text/html"),
    `browser gets ${html.status} ${html.type}`,
  );
  report(
    markdown.status === 200 && markdown.type.startsWith("text/markdown"),
    `agent gets ${markdown.status} ${markdown.type}`,
  );
  report(
    again.type.startsWith("text/html"),
    `browser still gets ${again.type} after an agent request`,
  );
  report(
    variesOnAccept(html.headers) && variesOnAccept(markdown.headers),
    "both responses send Vary: Accept",
  );
  report(
    file.status === 200 && file.body === markdown.body,
    `negotiated Markdown matches ${new URL(mirror).pathname} byte for byte`,
  );

  const tags = linkTags(html.body);
  const canonical = tags.find((tag) => tag.rel === "canonical")?.href;
  const alternate = tags.find(
    (tag) => tag.rel === "alternate" && tag.type === "text/markdown",
  )?.href;
  report(
    canonical !== undefined &&
      new URL(canonical).pathname === new URL(url).pathname,
    `canonical is ${canonical ?? "missing"}`,
  );
  report(
    alternate !== undefined && new URL(alternate, url).href === mirror,
    `HTML names ${alternate ?? "no"} Markdown alternate`,
  );

  const headings = htmlHeadings(html.body);
  const missingHeadings = markdownHeadings(markdown.body).filter(
    (heading) => !headings.some((text) => text.includes(heading)),
  );
  report(
    missingHeadings.length === 0,
    missingHeadings.length === 0
      ? "every Markdown heading is in the HTML"
      : `headings missing from the HTML: ${missingHeadings.join(" | ")}`,
  );

  const hrefs = htmlLinks(html.body);
  const missingLinks = markdownLinks(markdown.body).filter(
    (link) => !hrefs.has(link),
  );
  report(
    missingLinks.length === 0,
    missingLinks.length === 0
      ? "every root-relative Markdown link is in the HTML"
      : `links missing from the HTML: ${missingLinks.join(" ")}`,
  );

  const cache = [html, markdown, again].map((response) =>
    cacheInfo(response.headers),
  );
  if (cache.some(Boolean))
    console.log(`  cache: ${cache.map((info) => info || "-").join(" / ")}`);
}

console.log(
  failures === 0 ? "\nAll checks passed." : `\n${failures} check(s) failed.`,
);
process.exit(failures === 0 ? 0 : 1);

Run it against a few content pages:

node check-negotiation.mjs https://docs.acme.example /send-a-message /authentication

For each path, it requests the page as a browser, as an agent, and as a browser again, then fetches the .md file. It checks:

  • The formats. The browser gets a 200 HTML page both times, and the agent gets 200 Markdown.
  • The cache header. Both responses carry Vary: Accept.
  • The body. The negotiated Markdown is byte for byte the same as the .md file, so the two addresses can't drift apart.
  • The identity. The HTML's canonical URL is the page's own path, and its Markdown alternate points at the .md file.
  • The content. Every heading in the Markdown appears in the HTML, and so does every root-relative link.

To see a passing run before your own deploy is up, point it at useblume.dev, Blume's own docs, which deploy with cloudflare():

node check-negotiation.mjs https://useblume.dev /docs/quickstart
/docs/quickstart
  ok   browser gets 200 text/html
  ok   agent gets 200 text/markdown; charset=utf-8
  ok   browser still gets text/html after an agent request
  ok   both responses send Vary: Accept
  ok   negotiated Markdown matches /docs/quickstart.md byte for byte
  ok   canonical is https://useblume.dev/docs/quickstart
  ok   HTML names /docs/quickstart.md Markdown alternate
  ok   every Markdown heading is in the HTML
  ok   every root-relative Markdown link is in the HTML
  cache: cf-cache-status: HIT / cf-cache-status: HIT / cf-cache-status: HIT

All checks passed.

The script exits with status 1 when a check fails, so it can run in CI after each deploy. Leave a landing-page homepage out of the run: its Markdown is the llms.txt index, not a copy of the page, so the heading check fails there by design. Run it against a deployment rather than blume dev, too. The dev server sends Vary: Accept only on the Markdown response, since no cache sits in front of it, so that one check fails locally.

When the formats differ on purpose

The script checks that what the Markdown says is also on the page, not the reverse, so it won't notice what the Markdown leaves out. Two features can leave things out, so check pages that use them by hand:

  • Content in <Visibility for="web"> is left out of the Markdown, and content in <Visibility for="agents"> appears only there. Keep both to asides, like a note telling agents which page to read next, so the two formats still cover the same material.
  • A component of your own with no Markdown form stays as JSX in the .md file, so an agent reads its props, not what it renders. Give it a Markdown serializer in agents.markdownComponents.

Check what your CDN caches

Two formats at one URL only work if every cache between the client and your host keys its entries on Accept. Otherwise the first response a cache stores is what everyone gets: agents receive HTML, or browsers receive raw Markdown.

On the two hosts that negotiate, the host's own layer is covered:

  • Vercel's CDN includes the Accept header in its cache key by default, and Blume adds Vary: Accept to both formats of every negotiated URL.
  • On Cloudflare, Workers run in front of the cache, so the generated Worker chooses the file, the page or its .md, on every request.

The risk is a layer you add in front, like a second CDN or a proxy on another provider. Cloudflare's cache, for example, doesn't consider Vary values by default, except Accept-Encoding, so a rule that makes it cache your HTML pages stores whichever format arrived first. If you run a layer like that, configure it to vary on Accept, or keep it from caching pages.

The script's browser, agent, browser order already tests this, and its cache: line prints the cache status headers each response carried. Run it twice in a row, so the second run reads from whatever the first one warmed.

Troubleshooting

The agent gets HTML on the deployed site

Look for the Wired Accept: text/markdown negotiation line in the build log. If it's missing, the build isn't a Vercel or Cloudflare server build: check that deployment names vercel() or cloudflare() without output: "static". On a static host, Netlify, or node(), use the .md URLs instead. Pages you wrote as .astro files in pages/ have no Markdown mirror, so they always answer with HTML, except a landing-page homepage, which answers with llms.txt.

The Cloudflare build says it could not wire negotiation

Blume merges its routing rules into any run_worker_first list in your own wrangler.jsonc. When the combined list breaks Wrangler's limits of 100 rules of 100 characters each, Blume leaves the routing alone and warns, and the .md URLs keep working. Trim your own rules until the combined list fits, then build again.

Dev and production answer the same request differently

The dev server and the Cloudflare Worker compare quality values, so Accept: text/markdown;q=0.5, text/html gets HTML there. Vercel's routing rule matches any lowercase text/markdown entry without comparing quality values, so the same request gets Markdown on Vercel. Send text/markdown alone, or with the highest quality value, and every host answers the same way.

The canonical check fails

A missing canonical means Blume doesn't know the site URL: set site in the adapter, as in the Cloudflare config above. Some pages point elsewhere on purpose. A page marked noindex has no canonical, and a page in an archived version names the latest version's page as its canonical by default, so leave those out of the run.

Every check fails on a Vercel preview

Deployment Protection answers the script with a login page instead of your docs. Run it against production, or add your project's Protection Bypass for Automation secret to the headers in the script's get function, as an x-vercel-protection-bypass header.

Next step

Add an MCP server

A server build on Vercel or Cloudflare can also host an MCP server, so agents can search your docs and fetch pages as tools.

Read the MCP server guide

A step here not working for you? Report a broken step.

Keep going.More guides.

Upgrade your docs with Blume.

Install today and ship a production-grade docs site in minutes. Free and open source, forever.

npx blume init