Skip to content
Blume
Esc
↑↓navigate↵open⌘Jpreview
Guides

Agents

Rate-limit a public documentation AI assistant

A per-reader limit on your docs assistant that holds across serverless instances, a load test that shows the 429s and the reset, and a bot check for scripts.

By 10 min read

To stop one visitor from sending your docs assistant request after request, count requests per visitor at the server and refuse the extras. Blume does this by default. The assistant's /api/ask route lets each IP address ask 30 questions every 10 minutes. Past that, it answers 429 Too Many Requests with a Retry-After header, and the panel tells the reader to try again in a few minutes.

The catch is where the count lives. By default it's in the server's memory, which is exact on a single Node process but separate for every instance on Vercel, Netlify, and Cloudflare. In this guide you reproduce the limit locally with a small load-test script, move the count to a shared store so it holds across instances, test the deployed limit and its recovery, and add a bot check for scripts that rotate addresses.

If you run one Node server that sees each reader's own address, the default already holds and you only need to pick the numbers. If your assistant answers from your own backend through the endpoint option, Blume generates no route, so rate limiting is that backend's job. And a rate limit is neither a spending cap nor a sign-in: what a rate limit doesn't do covers what to pair it with.

How the limit works

The limit check is the first thing the route does, before it reads the body, checks a bot token, or calls the model. So every request counts, including ones the route then refuses. The count is kept per site hostname, per route, and per IP address, which means the assistant, the API playground proxy, and server-side search each have their own budget.

Two cases are always let through, so readers are never locked out by accident: a request whose address the host can't tell, and one the shared store fails to count. The second is logged. A static site has no server routes, so there's nothing to limit.

These are the responses you'll see while testing:

StatusBody starts withMeaning
429Too many requests: try again inOver the limit. Retry-After gives the seconds left.
400Invalid request:Counted, then refused for a malformed body.
403Verification failed:Counted, then refused by the bot check.
503The assistant is not configured:A provider key or bot-check secret is missing.

Reproduce the limit locally

Start with numbers small enough to hit in seconds. Turn the assistant on, name a host adapter (the assistant needs server output), and set a limit of 5 requests a minute. If you haven't set up the assistant yet, add it first.

import { defineConfig } from "blume";
import { vercel } from "blume/deploy";
import { memory } from "blume/ratelimit";

export default defineConfig({
  title: "Acme Docs",
  deployment: vercel(),
  ai: {
    assistant: { enabled: true },
  },
  // Test numbers: 5 requests per reader per 60 seconds.
  rateLimit: memory({ requests: 5, window: 60 }),
});

Next, save the load-test script. It sends requests with an empty message list. The route counts each one, then refuses it with a 400 before the model runs, so the test spends no tokens and works without an API key. It prints each status it got, waits out the window, and sends one more request to check the limit resets.

// Sends requests the assistant route counts but never answers, prints the
// statuses that came back, then waits out the window and checks it recovers.
//
//   node ask-load-test.mjs <url> [total] [concurrency]
const [
  url = "http://localhost:4321/api/ask",
  total = "40",
  concurrency = "1",
] = process.argv.slice(2);

const headers = { "content-type": "application/json" };
const bypass = process.env.VERCEL_AUTOMATION_BYPASS_SECRET;
if (bypass) {
  headers["x-vercel-protection-bypass"] = bypass;
}
// No messages: the route counts the request against the limit, then
// refuses it with a 400 before it calls the model.
const body = JSON.stringify({ messages: [] });

const send = async () => {
  const response = await fetch(url, { body, headers, method: "POST" });
  const text = await response.text();
  return {
    retryAfter: Number(response.headers.get("retry-after") ?? 0),
    status: response.status,
    text: text.slice(0, 80),
  };
};

const counts = new Map();
const samples = new Map();
let retryAfter = 0;
let started = 0;
const worker = async () => {
  while (started < Number(total)) {
    started += 1;
    const result = await send();
    counts.set(result.status, (counts.get(result.status) ?? 0) + 1);
    if (!samples.has(result.status)) {
      samples.set(result.status, result.text);
    }
    if (result.status === 429) {
      retryAfter = Math.max(retryAfter, result.retryAfter);
    }
  }
};
await Promise.all(Array.from({ length: Number(concurrency) }, worker));

for (const [status, count] of [...counts].sort((a, b) => a[0] - b[0])) {
  console.log(status + "  x" + count + "  " + samples.get(status));
}
const limited = counts.get(429) ?? 0;
console.log("Got past the limit: " + (Number(total) - limited) + " of " + total);

if (limited > 0) {
  console.log("Waiting " + (retryAfter + 1) + "s for the window to reset...");
  await new Promise((resolve) => setTimeout(resolve, (retryAfter + 1) * 1000));
  const after = await send();
  console.log("After the wait: " + after.status + "  " + after.text);
}

Start the dev server, and in a second terminal, send 8 requests:

npx blume dev
node scripts/ask-load-test.mjs http://localhost:4321/api/ask 8

The first five reach validation, and the rest are limited:

400  x5  Invalid request: send 1-40 user/assistant messages with string content.
429  x3  Too many requests: try again in 60 seconds.
Got past the limit: 5 of 8
Waiting 61s for the window to reset...
After the wait: 400  Invalid request: send 1-40 user/assistant messages with string content.

Before the wait ends, open the assistant in the browser on the same machine and ask something. The browser shares the script's address, so the panel answers "You've asked a lot of questions. Try again in a few minutes." instead of calling the model. After the wait, it answers normally, once a provider key is set.

With memory() and upstash(), the window opens with a reader's first request and resets when it ends; it doesn't slide. Restarting the dev server clears a memory count.

Share the count across serverless instances

On Vercel, Netlify, and Cloudflare, your route runs on many short-lived instances, and memory() gives each its own count. It still stops a fast burst, but a reader whose requests land on several instances can get more than the limit. For an exact limit, keep the count in a store every instance shares:

AdapterCount kept inLimit holds
memory()The server's memory (the default)Exactly on one process; per instance on serverless hosts
upstash()Upstash RedisExactly, on every host
cloudflare()Workers rate limitingPer Cloudflare location

Upstash on Vercel

Create an Upstash Redis database, either from the Vercel Marketplace or in the Upstash console. Blume talks to it over Upstash's REST API, so there's no package to install. It reads two variables: UPSTASH_REDIS_REST_URL and UPSTASH_REDIS_REST_TOKEN.

The Marketplace connects the database to your project as KV_REST_API_URL and KV_REST_API_TOKEN, and Blume doesn't read those names. In your project's environment variables, add UPSTASH_REDIS_REST_URL and UPSTASH_REDIS_REST_TOKEN with the same values. Then switch the adapter, back at the default numbers:

import { defineConfig } from "blume";
import { vercel } from "blume/deploy";
import { upstash } from "blume/ratelimit";

export default defineConfig({
  title: "Acme Docs",
  deployment: vercel(),
  ai: {
    assistant: { enabled: true },
  },
  rateLimit: upstash({ requests: 30, window: 600 }),
});

To use the same database locally, put both variables in .env.local. Then run npx blume doctor: its summary should read Rate limiting: upstash, with no BLUME_MISSING_SECRET warning for either variable. Without them, the routes quietly count in memory, so this check matters. Variables only reach new deployments, so push or redeploy after adding them.

Cloudflare Workers

On a site deployed with deployment: cloudflare(), count with Cloudflare's own rate limiting instead. Blume declares the binding in the Worker's config when it builds, so there's no store to create:

import { defineConfig } from "blume";
import { cloudflare } from "blume/deploy";
import { cloudflare as cloudflareRateLimit } from "blume/ratelimit";

export default defineConfig({
  title: "Acme Docs",
  deployment: cloudflare(),
  ai: {
    assistant: { enabled: true },
  },
  rateLimit: cloudflareRateLimit({ requests: 10, window: 60 }),
});

The window must be 10 or 60 seconds. Cloudflare counts per location and describes the counter as permissive and eventually consistent, so treat the number as approximate. blume build logs "Declared the Workers rate limiting binding" when it's in place. The Cloudflare Workers guide covers the deploy.

Test the deployed limit and recovery

A small test against memory() on a serverless host can pass by luck, when one instance happens to serve every request. So test with the shared adapter in place, and send requests in parallel so they spread across instances. Replace docs.acme.example with your deployment's domain:

node scripts/ask-load-test.mjs https://docs.acme.example/api/ask 40 8

With upstash() at 30 requests, the script should print Got past the limit: 30 of 40. Anything above 30 means the count isn't shared. For comparison, with two instances behind a round-robin balancer and a limit of 5, memory() let 10 of 20 requests through and upstash() let exactly 5. The script then waits out the window, up to 10 minutes at these numbers, and its last request should get a 400 again.

If the deployment uses Vercel Deployment Protection, the script's requests stop at the login page. Export the project's automation bypass secret as VERCEL_AUTOMATION_BYPASS_SECRET first, and the script sends it in the x-vercel-protection-bypass header.

You can also look at the counters in the database:

curl -s -X POST "$UPSTASH_REDIS_REST_URL" \
  -H "Authorization: Bearer $UPSTASH_REDIS_REST_TOKEN" \
  -d '["KEYS", "blume:*"]'

Each key is blume:, then the hostname, the route, and the address, like blume:docs.acme.example:ask:203.0.113.7, and it expires when its window ends. Give the counters a database of their own, since KEYS scans everything in it.

The test uses up your own address's budget, so the production assistant refuses you until the window ends. To see whether real readers hit the limit, configure an analytics provider: a refused question is reported as an ask_error event with status 429. If those show up for ordinary use, raise requests.

Add a bot check

Counting by address stops one machine. A script that rotates through many addresses gets a fresh budget on each. To stop it, add a bot check from blume/captcha. Before each question, the panel gets a token from the check, and the route verifies it with the provider before the model runs. Most readers never see a challenge.

In the Cloudflare dashboard, add a Turnstile widget, list your docs domain under hostname management, and choose the Managed or Invisible mode. For local testing, Cloudflare publishes an invisible test site key that always passes:

import { defineConfig } from "blume";
import { turnstile } from "blume/captcha";
import { vercel } from "blume/deploy";
import { upstash } from "blume/ratelimit";

export default defineConfig({
  title: "Acme Docs",
  deployment: vercel(),
  ai: {
    assistant: {
      enabled: true,
      // Cloudflare's always-pass test key. Use your widget's site key to deploy.
      captcha: turnstile({ siteKey: "1x00000000000000000000BB" }),
    },
  },
  rateLimit: upstash({ requests: 30, window: 600 }),
});
TURNSTILE_SECRET_KEY=1x0000000000000000000000000000000AA

The site key is public and safe to commit. Set the widget's real secret key as TURNSTILE_SECRET_KEY in your host's environment. Until it's set, blume build warns and the assistant answers that it isn't configured. To use hCaptcha instead, swap in hcaptcha() from the same module, with HCAPTCHA_SECRET_KEY.

A failed check answers 403, and the panel says it couldn't check that the reader is human. The rate limit still runs first, so failed checks count too. The load-test script keeps working, since its requests are refused for their empty body before the check.

What a rate limit doesn't do

It isn't a spending cap. Blume bounds each request: at most 40 messages and 24,000 characters of conversation, a 64 KB body, and five model steps per question. The limit then bounds each address. Nothing bounds the total across many addresses, so set a budget or usage limit with your model provider as well. The pricing page covers what the assistant costs to run.

It isn't authentication. The route has to be callable without signing in, because the in-page panel calls it. If only your customers should reach the assistant, protect the whole site with your host's access protection, or point endpoint at a backend that checks sessions itself.

It can't tell people apart on a shared address. An office network or a mobile carrier can put many readers behind one IP address, and they share one count. If your readers mostly sit behind one network, raise requests.

Troubleshooting

Every request gets 403 "Cross-site POST form submissions are forbidden"

Astro's cross-site check refuses a POST with no content type, or a form-like one such as text/plain, before the route runs. Send content-type: application/json, as the script does. These refusals don't count against the limit.

The deployed limit lets through more than you set

Check that the adapter really is shared. Your function logs say "Rate limiting counts in memory: set UPSTASH_REDIS_REST_URL and UPSTASH_REDIS_REST_TOKEN to share the count through Upstash." when the variables didn't reach the deployment, and "Rate limiting failed; letting the request through" when the store can't be reached. Also check that rateLimit isn't set to false.

The build warns BLUME_MISSING_SECRET for UPSTASH_REDIS_REST_URL

The Marketplace names don't count. Add the two UPSTASH_ variables with the same values as KV_REST_API_URL and KV_REST_API_TOKEN, then redeploy.

Everyone hits the limit at once on a self-hosted server

On deployment: node() behind a reverse proxy or load balancer, the server can see the proxy's address on every request. Astro only trusts X-Forwarded-For when the request's host matches its security.allowedDomains setting, which Blume doesn't set, so every reader shares one count. Limit the route in the proxy instead and set rateLimit: false.

One reader stays limited long after the window

With upstash(), check that reader's key. Blume opens a window and counts the request in separate commands, so a key that expires between them comes back without an expiry, and that reader's count never resets. A TTL of -1 means this happened; delete the key to reset it:

curl -s -X POST "$UPSTASH_REDIS_REST_URL" \
  -H "Authorization: Bearer $UPSTASH_REDIS_REST_TOKEN" \
  -d '["TTL", "blume:docs.acme.example:ask:203.0.113.7"]'

curl -s -X POST "$UPSTASH_REDIS_REST_URL" \
  -H "Authorization: Bearer $UPSTASH_REDIS_REST_TOKEN" \
  -d '["DEL", "blume:docs.acme.example:ask:203.0.113.7"]'

Readers see "We couldn't check that you're human"

The bot check failed in the browser or at the route. Make sure the widget's hostname list includes your docs domain, and that the site key and secret key come from the same widget. Cloudflare's test tokens only pass with the test secret, and production secrets reject them.

Next step

Check which limiter runs

Run it after switching to upstash(): the summary names the rate limiter, and a missing Upstash variable shows up as a warning.

npx blume doctor
Read the rate limiting docs

A step here not working for you? Report a broken step.

Keep going.More guides.

Upgrade your docs with Blume.

Install today and ship a production-grade docs site in minutes. Free and open source, forever.

npx blume init