Skip to main content

TLS

nginx terminates TLS on :443. A separate caddy container is used purely as an ACME client to obtain and auto-renew a free Let's Encrypt certificate once you have a real domain — it never serves application traffic itself.

Why two containers instead of one​

Caddy's headline feature is automatic HTTPS, but this stack's routing (the cookie↔Authorization-header translation in nginx/auth.js, the sandbox's deliberately-unproxied path, per-path proxy_read_timeouts for SSE, etc.) is already implemented in nginx/njs. Replacing nginx would mean re-implementing all of that. Instead:

  • nginx stays the single public entry point on :80/:443 and keeps 100% of its existing routing, unchanged.
  • Caddy only ever talks to Let's Encrypt. It isn't published to the host and is unreachable from outside the reelbolt docker network.

nginx reads the certificate Caddy obtains directly off a shared volume (caddy_certs, mounted read-only into nginx). There is no copying, syncing, or cross-container signaling: renewals just update the file in place, and nginx's own background loop (in nginx/docker-entrypoint.sh) runs nginx -s reload every 12 hours to pick up the new file — a no-op when nothing changed.

Local development (no domain)​

Nothing to configure. docker compose up generates a self-signed certificate the first time nginx starts (nginx/docker-entrypoint.sh) and serves it on :443 alongside the existing :80. Browsers will show a certificate warning for the self-signed cert — that's expected, not a bug. FORCE_HTTPS defaults to false, so http://localhost keeps working exactly as before TLS support was added.

Going live with a real domain​

  1. Point the domain's DNS A/AAAA record at this host's public IP, and make sure :80 and :443 are actually reachable from the internet (this is required for Let's Encrypt's HTTP-01 challenge, which always connects on port 80 first).
  2. Set DOMAIN and ACME_EMAIL in .env.
  3. docker compose --profile tls up -d caddy — Caddy is opt-in via the tls compose profile so it never even starts (let alone crash-loops on a missing DOMAIN) for anyone who hasn't gone through this setup.
  4. docker compose up -d --force-recreate nginx — nginx only re-evaluates which certificate to load at container start (nginx/docker-entrypoint.sh runs once, at boot), so switching from "no domain" to "has a domain" needs this one manual recreate. After this, renewals are automatic — no further steps, ever.
  5. Confirm https://<DOMAIN> works and shows a real, trusted certificate.
  6. Only then, set FORCE_HTTPS=true in .env and restart nginx. This makes http:// redirect to https://. Flipping it on before step 5 is confirmed just replaces a working http:// dev flow with a broken or self-signed-warning https:// one.
  7. Set COOKIE_SECURE=true in .env and restart go-api. Do this only after https://<DOMAIN> is confirmed working — browsers silently drop Secure cookies sent over plain HTTP, so setting this too early breaks login with no obvious error.

How the pieces fit together​

  • nginx/Dockerfile — extends nginx:alpine with openssl (self-signed cert generation) and gettext (envsubst, used to render the HTTPS server block's certificate paths).
  • nginx/docker-entrypoint.sh — on every container start: picks a certificate (Caddy's, if DOMAIN is set and one has actually been issued; otherwise a self-signed one, generated once and reused), renders nginx/https-server.conf.template into /etc/nginx/conf.d/https-server.conf, copies either nginx/http-server-dev.conf or nginx/http-server-redirect.conf into /etc/nginx/conf.d/http-server.conf depending on FORCE_HTTPS, starts the periodic reload loop, then execs nginx.
  • nginx/locations.conf — the actual routing table (unchanged from before TLS support), shared via include by both the :80 (dev mode) and :443 server blocks so it only exists in one place.
  • nginx/http-server-dev.conf / nginx/http-server-redirect.conf — the two possible :80 server blocks (passthrough vs. redirect-to-https). Both also forward /.well-known/acme-challenge/ to the caddy container, which is how Let's Encrypt's HTTP-01 validation reaches Caddy even though Caddy itself is never exposed to the internet.
  • caddy/Caddyfile — a minimal site block for {$DOMAIN} that does nothing but respond 200 (Caddy manages a certificate for any domain used as a site address in its config, automatically, on startup).
  • caddy/docker-entrypoint.sh — refuses to start with a clear error if DOMAIN/ACME_EMAIL aren't set, instead of Caddy's own less obvious error from an empty {$DOMAIN} site address.

Caddy's on-disk certificate storage layout is a stable, documented convention: /data/caddy/certificates/acme-v02.api.letsencrypt.org-directory/<domain>/<domain>.crt (and .key) for the production Let's Encrypt CA. nginx/docker-entrypoint.sh reads this path directly (mounted read-only at /caddy-data) — there's nothing else keeping the two containers in sync.

In the cloud (Cloudflare in front of nginx)​

Production and staging (plan saas-launch, D19) keep nginx as the in-cluster router, since the cookie to Authorization translation in nginx/auth.js and the per-path SSE timeouts live there, and put the Cloudflare proxy (Free plan) in front for TLS, WAF and DDoS. Set Cloudflare SSL mode to Full (strict).

Two ways to give the origin a certificate, and both are in use:

  • Caddy + Let's Encrypt, the same mechanism as self-host: enable the compose tls profile and set DOMAIN/ACME_EMAIL. HTTP-01 works through the Cloudflare proxy because Cloudflare forwards the challenge path to the origin — this is what test.reelbolt.ai runs. See "Renewals" below for the one non-obvious requirement.
  • A Cloudflare Origin CA certificate for <domain> and *.<domain>, pointed at with TLS_CERT_FILE/TLS_KEY_FILE (mounted files; this overrides the Caddy and self-signed choice in docker-entrypoint.sh). Pick this when Cloudflare's SSL mode is Full (strict) and you would rather Cloudflare itself vouched for the origin, or when you need a wildcard. A wildcard matters because the runner. hostname below is a second name on the same listener; a single-name Let's Encrypt certificate does not cover it.

nginx behaviour is controlled by environment variables, all defaulting to the single-host values so local compose behaves as before:

VariableDefaultMeaning
UPSTREAM_GO_API, UPSTREAM_INFERENCE_API, UPSTREAM_WORKFLOW_ENGINE, UPSTREAM_WEB, UPSTREAM_SITE, UPSTREAM_DOCS, UPSTREAM_CADDYthe compose service URLs (http://workflow-engine:8080, ...)Upstream base URLs, substituted into /etc/nginx/conf.d/common.inc at start. The cloud control VM points UPSTREAM_WORKFLOW_ENGINE at http://$ENGINE_PRIVATE_IP:8080.
REAL_IP_FROM_CLOUDFLAREfalsetrue installs nginx/cloudflare-realip.conf: set_real_ip_from for Cloudflare's published ranges and real_ip_header CF-Connecting-IP. Needed for per-IP rate limits (C5) and logs behind the proxy. Only connections from those ranges are trusted, so a direct hit on the origin cannot spoof an address. Refresh the ranges from cloudflare.com/ips-v4 and /ips-v6 when they change. Leave off locally.
CLIENT_MAX_BODY_SIZE100MRequest body limit for the main server blocks.
UPSTREAM_RUNNER_GATEWAY, RUNNER_HOSTNAMEempty, runner.$DOMAINSee below.
TLS_CERT_FILE, TLS_KEY_FILEemptyOperator-supplied certificate, e.g. the Origin CA cert.

runner.​

docs/runner-protocol.md has runners dial wss://runner.<domain>/v1/connect. When UPSTREAM_RUNNER_GATEWAY is set, the entrypoint renders nginx/runner-server.conf.template: a separate server block for RUNNER_HOSTNAME that proxies only /v1/connect to the gateway with WebSocket upgrade headers, without the cookie translation (runners authenticate with a signed Ed25519 hello inside the protocol) and with any client Authorization header cleared, proxy_buffering off, and 120 s read/send timeouts. The protocol heartbeats every 15 s and declares a runner dead after 45 s, so the proxy never cuts a healthy idle socket; Cloudflare's own WebSocket idle limit (100 s) is also above 45 s. Everything else on that hostname is a 404. The gateway does not exist yet (decision D5a), so the variable is unset in the cloud override and no block is rendered (an empty file is included). TODO(runner-gateway WP): add the runner-gateway service and set RUNNER_GATEWAY_UPSTREAM in the control VM's .env, create the proxied runner DNS record, and run the one-hour WebSocket soak from F15's acceptance test. If UPSTREAM_RUNNER_GATEWAY is set with neither RUNNER_HOSTNAME nor DOMAIN, nginx refuses to start.

Renewals​

This deployment obtains its certificate with Caddy over ACME HTTP-01, which works through the Cloudflare proxy: Cloudflare forwards /.well-known/acme-challenge/ to the origin, where nginx proxies it to the caddy container. Two things are worth knowing before changing any of it.

Cloudflare SSL mode must be Full or Full (strict) — never Flexible. In the Full modes Cloudflare reaches the origin on :443; in Flexible it fetches :80. With FORCE_HTTPS=true that is an infinite redirect loop, because nginx answers every :80 request with a 301 to https:// that Cloudflare then re-requests over plain HTTP. Check which mode is in use by sampling the established sockets on the origin while driving traffic through the edge:

ss -Htn state established | awk '{print $4}' | grep -oE ':(80|443)$' | sort | uniq -c

:443 means Full or Full (strict), and FORCE_HTTPS is then safe. Full (strict) is the right target: it makes Cloudflare validate the origin certificate, so an attacker cannot impersonate the origin to the edge.

The Host header on the challenge location is load-bearing. Caddy matches an incoming challenge to the site it manages by hostname, so it only answers while it sees Host: <DOMAIN>. nginx's default for proxy_pass is Host: $proxy_host — here the upstream name, the bare word caddy, which Caddy does not manage. The request then falls through to Caddy's own HTTP→HTTPS redirect (a 308 to https://caddy/.well-known/...) and Let's Encrypt refuses that URL outright: invalid host: must end in IANA registered TLD.

This failure is invisible until renewal. A certificate obtained before FORCE_HTTPS was switched on keeps working, so the site looks healthy for ~60 days and then the certificate simply expires. nginx/http-server-dev.conf has always forwarded Host; nginx/http-server-redirect.conf did not, and only the redirect block is used once FORCE_HTTPS=true. It now forwards Host and the forwarding headers, matching the dev block.

Prove renewal rather than assume it. A cached ACME authorization will happily hide the fault, so clear the account as well as the certificate, and keep a copy first:

docker exec reelbolt-caddy-1 sh -c 'mv /data/caddy/certificates /data/caddy/certificates.bak; rm -rf /data/caddy/acme'
docker restart reelbolt-caddy-1
docker logs reelbolt-caddy-1 2>&1 | grep -E 'served key authentication|authz_status|certificate obtained|challenge failed'

served key authentication followed by certificate obtained successfully is a pass. Restore certificates.bak if it is not. nginx reads the certificate only at start or on reload, so after a re-issue run docker exec reelbolt-nginx-1 nginx -s reload or recreate the container — and note that if nginx starts while no Caddy certificate exists yet it silently falls back to the self-signed one, and keeps serving that until it is restarted.

Upload size​

Cloudflare Free rejects request bodies over 100 MB, and nginx's own limit is 100M, while F15's acceptance asks for a 500 MB upload. Options considered:

  1. Presigned direct-to-R2 browser uploads (the browser PUTs to R2, nginx and Cloudflare's proxy never see the bytes). Cleanest and also removes upload load from the control VM, but it needs an Inference API endpoint that issues presigned URLs, an R2 CORS policy, and a client change. Recommended, as a future work package.
  2. An unproxied (DNS-only) upload.<domain> hostname straight to nginx. Cheap in nginx, but it exposes the origin IP, bypasses WAF/DDoS for that name, and needs the auth cookie to be valid on a second host (Domain= on reelbolt_token, which does not exist today) plus CORS from the app host.
  3. Pay for a Cloudflare plan with a larger cap.

Decision for the beta: keep the 100 MB cap (documented limit; larger sources wait for option 1). Only the config side ships now: CLIENT_MAX_BODY_SIZE lets an operator raise the nginx limit if the proxy cap is later lifted or an unproxied host is added. No upload. server block is implemented, because without the cookie-domain and CORS work it would not authenticate.

SSH​

Deploy SSH from GitHub-hosted runners and its firewall and sshd hardening are described in docs/ci.md ("SSH reachability for deploys").