Skip to content
Blume
Esc
↑↓navigate↵open⌘Jpreview
Guides

Hosting

Host static documentation on S3 and CloudFront

A private S3 bucket behind CloudFront that serves every page at its clean URL, Markdown and JSON with the right types, real 404s, and a cache each deploy clears.

By 11 min read

To host Blume docs in an AWS account, build them to static files with blume build, upload dist/ to a private S3 bucket, and serve the bucket through CloudFront with origin access control. Two details decide whether it works: a small CloudFront Function that maps page URLs like /quickstart to the index.html Blume writes for them, and uploads that give each file the right media type and cache lifetime.

You end up with one CloudFormation stack (the bucket, the distribution, the function, and the bucket policy), a deploy script that builds, uploads, and clears the cache, and a check script that tests pages, Markdown, JSON, 404s, and caching on the live site. The docs get their own subdomain, docs.acme.example in the examples.

Choose this when the docs have to live in your AWS account. If they don't, Vercel, Netlify, and Cloudflare serve the same dist/ with less setup, because Blume detects the site URL there and writes the header and redirect files they read. Everything here is static, so features that answer requests at runtime, like the MCP server, need a server.

Choose the S3 endpoint

CloudFront can read a bucket through two S3 endpoints. They differ in who can read the bucket and in how a URL finds its file:

REST endpoint (this guide)Website endpoint
Bucket accessPrivate; CloudFront signs each request with origin access controlObjects must be public; origin access control isn't supported
CloudFront to S3HTTPSHTTP only
A request for /quickstartLooks up the key quickstart exactly, so a function adds /index.htmlAnswers 302 to /quickstart/, then serves quickstart/index.html
A missing file403, or 404 when CloudFront may list the bucketThe bucket's error document
Redirect rulesNonePer object and per bucket

Blume writes each page as <route>/index.html and links to it without a trailing slash, which is also the canonical URL in the page's head. The website endpoint answers every one of those links with a redirect to the slashed form, and it needs a public bucket. The REST endpoint keeps the bucket private and serves the canonical URL directly once a function rewrites it, so that's the setup here.

Set the site URL and build

S3 and CloudFront don't tell the build where the site lives, so set it yourself. Blume needs the origin for the sitemap, canonical URLs, and Open Graph images:

import { defineConfig } from "blume";

export default defineConfig({
  title: "Acme",
  deployment: {
    site: "https://docs.acme.example",
  },
});

Leave deployment as a plain object: naming a host adapter switches the build to server output. Then build and preview it the way a static host serves it:

npx blume build
npx blume preview

Here's what lands in dist/ and matters on S3:

  • index.html and a <route>/index.html for every page, plus 404.html.
  • A Markdown mirror of each page at <route>.md and <route>.mdx, with the home page's at index.md, and llms.txt, llms-full.txt, robots.txt, and sitemap.xml.
  • JSON: the search index (blume-search.json), the page index at api/docs/pages.json with a document per page, and openapi.json.
  • Discovery files in .well-known/. One of them, api-catalog, has no file extension.
  • Fingerprinted CSS and JavaScript in _astro/.
  • Config for other hosts: _headers, vercel.json, and, when you have redirects, _redirects and blume-redirects.json. CloudFront reads none of them, so the deploy script applies their rules itself and leaves them out of the bucket.

Request a certificate

CloudFront only uses ACM certificates from us-east-1, whatever region your bucket is in. Request one for the docs domain, then print the DNS record ACM needs to validate it:

CERT_ARN="$(aws acm request-certificate --region us-east-1 \
  --domain-name docs.acme.example --validation-method DNS \
  --query CertificateArn --output text)"

aws acm describe-certificate --region us-east-1 --certificate-arn "$CERT_ARN" \
  --query "Certificate.DomainValidationOptions[0].ResourceRecord"

Add that CNAME record at your DNS provider. The certificate is issued once the record resolves, and the stack in the next step waits for it.

Create the bucket and distribution

This template creates everything that serves the site:

AWSTemplateFormatVersion: "2010-09-09"
Description: Static Blume docs in a private S3 bucket behind CloudFront

Parameters:
  DomainName:
    Type: String
    Description: The docs domain, like docs.acme.example.
  CertificateArn:
    Type: String
    Description: An ACM certificate for DomainName, requested in us-east-1.
    AllowedPattern: "^arn:aws[a-z-]*:acm:us-east-1:[0-9]{12}:certificate/.+$"

Resources:
  SiteBucket:
    Type: AWS::S3::Bucket
    Properties:
      PublicAccessBlockConfiguration:
        BlockPublicAcls: true
        BlockPublicPolicy: true
        IgnorePublicAcls: true
        RestrictPublicBuckets: true
      OwnershipControls:
        Rules:
          - ObjectOwnership: BucketOwnerEnforced

  OriginAccessControl:
    Type: AWS::CloudFront::OriginAccessControl
    Properties:
      OriginAccessControlConfig:
        Name: !Sub "${AWS::StackName}-oac"
        OriginAccessControlOriginType: s3
        SigningBehavior: always
        SigningProtocol: sigv4

  DirectoryIndexFunction:
    Type: AWS::CloudFront::Function
    Properties:
      Name: !Sub "${AWS::StackName}-directory-index"
      AutoPublish: true
      FunctionConfig:
        Comment: Serve /page from /page/index.html
        Runtime: cloudfront-js-2.0
      FunctionCode: |
        function handler(event) {
          var request = event.request;
          var uri = request.uri;

          // "/" and "/guides/" are folders: serve their index.html.
          if (uri.endsWith("/")) {
            request.uri = uri + "index.html";
            return request;
          }

          // "/guides/setup" has no extension in its last segment, so it's a
          // page: serve /guides/setup/index.html. "/guides/setup.md" and
          // "/_astro/app.js" are files and pass through unchanged.
          var last = uri.slice(uri.lastIndexOf("/") + 1);
          if (last.indexOf(".") === -1) {
            request.uri = uri + "/index.html";
          }
          return request;
        }

  Distribution:
    Type: AWS::CloudFront::Distribution
    Properties:
      DistributionConfig:
        Enabled: true
        HttpVersion: http2and3
        Aliases:
          - !Ref DomainName
        ViewerCertificate:
          AcmCertificateArn: !Ref CertificateArn
          SslSupportMethod: sni-only
          MinimumProtocolVersion: TLSv1.2_2021
        Origins:
          - Id: site
            DomainName: !GetAtt SiteBucket.RegionalDomainName
            OriginAccessControlId: !GetAtt OriginAccessControl.Id
            S3OriginConfig:
              OriginAccessIdentity: ""
        DefaultCacheBehavior:
          TargetOriginId: site
          ViewerProtocolPolicy: redirect-to-https
          Compress: true
          # Managed policy CachingOptimized: honors each file's Cache-Control.
          CachePolicyId: 658327ea-f89d-4fab-a63d-7e88639e58f6
          FunctionAssociations:
            - EventType: viewer-request
              FunctionARN: !GetAtt DirectoryIndexFunction.FunctionMetadata.FunctionARN
        CacheBehaviors:
          # Discovery files: served exactly as named, readable cross-origin.
          - PathPattern: /.well-known/*
            TargetOriginId: site
            ViewerProtocolPolicy: redirect-to-https
            Compress: true
            CachePolicyId: 658327ea-f89d-4fab-a63d-7e88639e58f6
            # Managed policy SimpleCORS: Access-Control-Allow-Origin: *
            ResponseHeadersPolicyId: 60669652-455b-4ae9-85a4-c4c02393f86c
        CustomErrorResponses:
          - ErrorCode: 404
            ResponseCode: 404
            ResponsePagePath: /404.html

  SiteBucketPolicy:
    Type: AWS::S3::BucketPolicy
    Properties:
      Bucket: !Ref SiteBucket
      PolicyDocument:
        Version: "2012-10-17"
        Statement:
          - Sid: AllowCloudFrontRead
            Effect: Allow
            Principal:
              Service: cloudfront.amazonaws.com
            Action: s3:GetObject
            Resource: !Sub "${SiteBucket.Arn}/*"
            Condition:
              StringEquals:
                AWS:SourceArn: !Sub "arn:${AWS::Partition}:cloudfront::${AWS::AccountId}:distribution/${Distribution}"
          # Lets S3 answer a missing file with 404 instead of 403.
          - Sid: AllowCloudFrontList
            Effect: Allow
            Principal:
              Service: cloudfront.amazonaws.com
            Action: s3:ListBucket
            Resource: !GetAtt SiteBucket.Arn
            Condition:
              StringEquals:
                AWS:SourceArn: !Sub "arn:${AWS::Partition}:cloudfront::${AWS::AccountId}:distribution/${Distribution}"

Outputs:
  BucketName:
    Value: !Ref SiteBucket
  DistributionId:
    Value: !Ref Distribution
  DistributionDomain:
    Value: !GetAtt Distribution.DomainName

What each part does:

  • SiteBucket blocks all public access and keeps the Bucket owner enforced ownership setting, which origin access control requires.
  • OriginAccessControl signs every request CloudFront sends to S3. Signing always (SigningBehavior: always) also keeps that connection on HTTPS.
  • DirectoryIndexFunction runs on each request before the cache lookup. A path whose last segment has no dot is a page, so /guides/setup becomes /guides/setup/index.html. Files like /guides/setup.md pass through unchanged.
  • Distribution uses the managed CachingOptimized policy: no cookies or query strings in the cache key, compression on, and each file cached for as long as its Cache-Control header says. The /.well-known/* behavior skips the function, so the extensionless discovery files are served as named, and adds Access-Control-Allow-Origin: * through the managed SimpleCORS policy, since agent registries read them cross-origin. A missing file serves 404.html with a 404 status.
  • SiteBucketPolicy lets only this distribution read the bucket. The s3:ListBucket statement makes S3 report a missing file as 404 rather than 403, so the error mapping above works. The function rewrites / to /index.html before a request reaches S3, so the permission never exposes a listing.

Wait for the certificate, deploy the stack in any region, and print its outputs:

aws acm wait certificate-validated --region us-east-1 --certificate-arn "$CERT_ARN"

aws cloudformation deploy --stack-name acme-docs --template-file docs-site.yaml \
  --parameter-overrides DomainName=docs.acme.example CertificateArn="$CERT_ARN"

aws cloudformation describe-stacks --stack-name acme-docs \
  --query "Stacks[0].Outputs" --output table

Point docs.acme.example at the DistributionDomain output: a CNAME record, or an alias record if the zone is in Route 53.

Upload with the right types and caching

S3 serves each file with the Content-Type and Cache-Control stored when it was uploaded, so the upload is where both get decided. This script does the whole deploy:

#!/usr/bin/env bash
# Build the docs, upload them to the stack's bucket, and refresh CloudFront.
set -euo pipefail

STACK="acme-docs"
output() {
  aws cloudformation describe-stacks --stack-name "$STACK" \
    --query "Stacks[0].Outputs[?OutputKey=='$1'].OutputValue" --output text
}
BUCKET="$(output BucketName)"
DISTRIBUTION="$(output DistributionId)"

# Fingerprinted assets never change. Everything else is revalidated by
# browsers on every visit and kept at the edge until the next deploy.
IMMUTABLE="public, max-age=31536000, immutable"
REVALIDATE="public, max-age=0, s-maxage=86400, must-revalidate"

# Files the CLI can't type correctly from their extension: pattern, type.
TYPED=(
  "*.md" "text/markdown; charset=utf-8"
  "*.mdx" "text/markdown; charset=utf-8"
  "*.txt" "text/plain; charset=utf-8"
  "*.tar.gz" "application/gzip"
  ".well-known/api-catalog"
  'application/linkset+json; profile="https://www.rfc-editor.org/info/rfc9727"'
)
# Never uploaded: assets go up separately, and the rest are the host config
# files Blume writes for other platforms.
EXCLUDES=(--exclude "_astro/*" --exclude "_headers" --exclude "_redirects"
  --exclude "vercel.json" --exclude "blume-redirects.json")

npx blume build

# 1. Assets first, so no new page points at a file that isn't there yet.
aws s3 sync dist/_astro "s3://$BUCKET/_astro" --cache-control "$IMMUTABLE"

# 2. The typed files, one pattern at a time.
for ((i = 0; i < ${#TYPED[@]}; i += 2)); do
  aws s3 sync dist "s3://$BUCKET" --delete --exclude "*" \
    --include "${TYPED[i]}" --content-type "${TYPED[i + 1]}" \
    --cache-control "$REVALIDATE"
  EXCLUDES+=(--exclude "${TYPED[i]}")
done

# 3. Everything else, deleting what the build no longer has.
aws s3 sync dist "s3://$BUCKET" --delete "${EXCLUDES[@]}" \
  --cache-control "$REVALIDATE"

# 4. Clear the edge caches, and wait so a check afterward sees this deploy.
INVALIDATION="$(aws cloudfront create-invalidation \
  --distribution-id "$DISTRIBUTION" --paths "/*" \
  --query Invalidation.Id --output text)"
aws cloudfront wait invalidation-completed \
  --distribution-id "$DISTRIBUTION" --id "$INVALIDATION"

Media types

The AWS CLI guesses each file's type from its extension. HTML, CSS, JavaScript, JSON, XML, and images come out right. Markdown doesn't: .mdx is never recognized, whether .md is depends on the Python the CLI runs on, and neither gets a charset. Without charset=utf-8, browsers read the raw Markdown as Windows-1252, so any non-ASCII text turns into mojibake. The extensionless api-catalog gets no guess at all.

The TYPED list sets those types explicitly. It mirrors the Content-Type rules in dist/_headers, so if your build's copy lists another one, add a pattern and type for it. See Content types for why Blume pins the charset. CloudFront compresses the types on its list, like HTML and JSON, but text/markdown isn't one of them, so the .md mirrors travel uncompressed.

Caching

Files in _astro/ get a new name whenever their content changes, so they're cached for a year as immutable. Everything else gets two lifetimes. max-age=0 with must-revalidate makes browsers check with CloudFront on each visit, which answers 304 Not Modified when nothing changed. s-maxage=86400 lets CloudFront keep its copy for a day, and the invalidation at the end of every deploy clears it as soon as the new files are in. A day is the ceiling if an invalidation is ever missed, not the normal delay.

Order and deletes

Assets go up first, so no new page points at a file that isn't there yet. --delete removes pages you deleted, but it skips _astro/ on purpose: a tab opened before the deploy may still load a script chunk it hasn't fetched yet, like the search dialog's. Prune old assets now and then, well after a deploy:

BUCKET="$(aws cloudformation describe-stacks --stack-name acme-docs \
  --query "Stacks[0].Outputs[?OutputKey=='BucketName'].OutputValue" --output text)"
aws s3 sync dist/_astro "s3://$BUCKET/_astro" --delete

To run the deploy from CI, the credentials need s3:ListBucket, s3:PutObject, and s3:DeleteObject on the bucket, cloudfront:CreateInvalidation and cloudfront:GetInvalidation on the distribution, and cloudformation:DescribeStacks for the outputs.

Check the live site

This script requests the URLs that break first on S3 and prints one line per check, then exits non-zero if any failed:

#!/usr/bin/env bash
# Check a deployed Blume site's routing, media types, and caching.
# Usage: ./check-site.sh https://docs.acme.example /guides/setup
set -uo pipefail

SITE="${1%/}"
PAGE="${2:?pass a page path, like /guides/setup}"
FAILED=0

# check PATH STATUS HEADER EXPECTED
# Passes when the response has STATUS and HEADER contains EXPECTED.
check() {
  local response status value
  response="$(curl -s -o /dev/null -D - -H "Origin: https://example.com" \
    "$SITE$1" | tr -d '\r')"
  status="$(printf '%s\n' "$response" | awk 'NR == 1 { print $2 }')"
  value="$(printf '%s\n' "$response" | grep -i "^$3:" | head -n 1 |
    cut -d ' ' -f 2-)"
  if [ "$status" = "$2" ] && [[ "$value" == *"$4"* ]]; then
    echo "ok    $1 -> $status, $3: $value"
  else
    echo "FAIL  $1 -> $status (want $2), $3: $value (want $4)"
    FAILED=1
  fi
}

check "$PAGE" 200 content-type "text/html"
check "$PAGE" 200 cache-control "max-age=0"
check "$PAGE/" 200 content-type "text/html"
check "$PAGE.md" 200 content-type "text/markdown; charset=utf-8"
check "/llms.txt" 200 content-type "text/plain; charset=utf-8"
check "/api/docs/pages.json" 200 content-type "application/json"
check "/sitemap.xml" 200 content-type "xml"
check "/.well-known/api-catalog" 200 content-type "application/linkset+json"
check "/.well-known/api-catalog" 200 access-control-allow-origin "*"
check "/no-such-page-$RANDOM" 404 content-type "text/html"

ASSET="$(curl -s "$SITE$PAGE" | grep -o '/_astro/[^"]*\.css' | head -n 1)"
check "$ASSET" 200 cache-control "immutable"

# The page was fetched above, so this request should come from the edge.
check "$PAGE" 200 x-cache "Hit from cloudfront"

exit "$FAILED"

Run it against the domain and a page a few levels deep:

./check-site.sh https://docs.acme.example /guides/setup

Before DNS resolves, pass the distribution's cloudfront.net domain instead. It checks that:

  • A deep link loads as HTML without redirecting, with a trailing slash too, and browsers are told to revalidate it.
  • The page's .md mirror and llms.txt carry charset=utf-8.
  • The JSON page index, the sitemap, and the API catalog have their media types, and the catalog answers cross-origin requests.
  • A URL with no page behind it returns 404 with the HTML 404 page.
  • The page's stylesheet is immutable, and a repeat request for the page comes from the edge. Each edge location fills its own cache, so if that last check reports a miss, run the script again.

Add it to the end of your deploy, so a bad upload fails the pipeline instead of waiting for a reader to find it.

What needs a server

Everything Blume prerenders works from S3: pages, the built-in search index, the Markdown mirrors, llms.txt, the JSON page index and page documents, Open Graph images, and the sitemap. What doesn't is anything that answers a request at runtime:

  • The assistant's generated endpoint. Point it at a backend you run with an external endpoint and the site stays static.
  • The MCP server.
  • The API playground's built-in proxy.
  • Server-side search providers, like Mixedbread.
  • The JSON API's live search at /api/docs/search.
  • Accept: text/markdown negotiation at a page's own URL. Agents fetch the .md URL instead, which each docs page links from its head.

A static build with any of the first four turned on stops with BLUME_SERVER_FEATURE_REQUIRED. To keep them, build with the node() adapter and run the server in a container, as Self-host documentation with Docker and Node.js shows. Static or server-rendered documentation helps you decide.

Troubleshooting

Every page returns 403

S3 refused CloudFront's request. Check that the bucket policy's AWS:SourceArn names this distribution, that the origin uses the origin access control, and that the deploy uploaded to the stack's bucket. If you removed the s3:ListBucket statement, a missing file also shows up as 403.

The home page loads, but other pages 404

Look for the page's index.html in the bucket, like guides/setup/index.html. If it's there, the function isn't running: check that it's associated with the default behavior as a viewer-request function, and that it's published, which AutoPublish: true does in the template.

Markdown shows garbled characters

The .md files went up without their type, usually from a plain aws s3 sync. Check with curl -sI https://docs.acme.example/guides/setup.md, then run deploy.sh again: every build rewrites every file, so the sync uploads them all with the right metadata.

A deploy doesn't show up

CloudFront is still serving its copy. Check that the invalidation ran and completed; the script waits for it. Browsers revalidate each visit, so readers see the new version as soon as the edge does.

Old URLs return 200 instead of a redirect

A static build serves each exact redirect as a small HTML page that sends the browser on, and pattern redirects get no page, so they 404 here. CloudFront reads neither _redirects nor vercel.json. For real status codes, return them from the viewer-request function, using the pairs in blume-redirects.json.

A page with a dot in its URL 404s

The function treats a last segment with a dot as a file, so a page built from v1.2.md at /v1.2 is never rewritten. Rename the file, like v1-2.md, and add a redirect from the old URL.

The sitemap is missing

The build summary shows no (set deployment.site): the build had no site URL. Set deployment.site as above and deploy again.

Next step

Check links before you upload

Run it at the top of deploy.sh, so a broken link stops the deploy before anything reaches the bucket.

npx blume validate
Read the validate docs

A step here not working for you? Report a broken step.

Keep going.More guides.

Upgrade your docs with Blume.

Install today and ship a production-grade docs site in minutes. Free and open source, forever.

npx blume init