TLS
nginx terminates TLS on :443. A separate caddy container is used purely as
an ACME client to obtain and auto-renew a free Let's Encrypt certificate once
you have a real domain — it never serves application traffic itself.
Why two containers instead of one
Caddy's headline feature is automatic HTTPS, but this stack's routing (the
cookie↔Authorization-header translation in nginx/auth.js, the sandbox's
deliberately-unproxied path, per-path proxy_read_timeouts for SSE, etc.) is
already implemented in nginx/njs. Replacing nginx would mean re-implementing
all of that. Instead:
- nginx stays the single public entry point on
:80/:443and keeps 100% of its existing routing, unchanged. - Caddy only ever talks to Let's Encrypt. It isn't published to the host
and is unreachable from outside the
reelboltdocker network.
nginx reads the certificate Caddy obtains directly off a shared volume
(caddy_certs, mounted read-only into nginx). There is no copying, syncing,
or cross-container signaling: renewals just update the file in place, and
nginx's own background loop (in nginx/docker-entrypoint.sh) runs
nginx -s reload every 12 hours to pick up the new file — a no-op when
nothing changed.
Local development (no domain)
Nothing to configure. docker compose up generates a self-signed certificate
the first time nginx starts (nginx/docker-entrypoint.sh) and serves it on
:443 alongside the existing :80. Browsers will show a certificate warning
for the self-signed cert — that's expected, not a bug. FORCE_HTTPS defaults
to false, so http://localhost keeps working exactly as before TLS support
was added.
Going live with a real domain
- Point the domain's DNS
A/AAAArecord at this host's public IP, and make sure:80and:443are actually reachable from the internet (this is required for Let's Encrypt's HTTP-01 challenge, which always connects on port 80 first). - Set
DOMAINandACME_EMAILin.env. docker compose --profile tls up -d caddy— Caddy is opt-in via thetlscompose profile so it never even starts (let alone crash-loops on a missingDOMAIN) for anyone who hasn't gone through this setup.docker compose up -d --force-recreate nginx— nginx only re-evaluates which certificate to load at container start (nginx/docker-entrypoint.shruns once, at boot), so switching from "no domain" to "has a domain" needs this one manual recreate. After this, renewals are automatic — no further steps, ever.- Confirm
https://<DOMAIN>works and shows a real, trusted certificate. - Only then, set
FORCE_HTTPS=truein.envand restart nginx. This makeshttp://redirect tohttps://. Flipping it on before step 5 is confirmed just replaces a workinghttp://dev flow with a broken or self-signed-warninghttps://one. - Set
COOKIE_SECURE=truein.envand restartgo-api. Do this only afterhttps://<DOMAIN>is confirmed working — browsers silently dropSecurecookies sent over plain HTTP, so setting this too early breaks login with no obvious error.
How the pieces fit together
nginx/Dockerfile— extendsnginx:alpinewithopenssl(self-signed cert generation) andgettext(envsubst, used to render the HTTPS server block's certificate paths).nginx/docker-entrypoint.sh— on every container start: picks a certificate (Caddy's, ifDOMAINis set and one has actually been issued; otherwise a self-signed one, generated once and reused), rendersnginx/https-server.conf.templateinto/etc/nginx/conf.d/https-server.conf, copies eithernginx/http-server-dev.confornginx/http-server-redirect.confinto/etc/nginx/conf.d/http-server.confdepending onFORCE_HTTPS, starts the periodic reload loop, then execs nginx.nginx/locations.conf— the actual routing table (unchanged from before TLS support), shared viaincludeby both the:80(dev mode) and:443server blocks so it only exists in one place.nginx/http-server-dev.conf/nginx/http-server-redirect.conf— the two possible:80server blocks (passthrough vs. redirect-to-https). Both also forward/.well-known/acme-challenge/to thecaddycontainer, which is how Let's Encrypt's HTTP-01 validation reaches Caddy even though Caddy itself is never exposed to the internet.caddy/Caddyfile— a minimal site block for{$DOMAIN}that does nothing but respond 200 (Caddy manages a certificate for any domain used as a site address in its config, automatically, on startup).caddy/docker-entrypoint.sh— refuses to start with a clear error ifDOMAIN/ACME_EMAILaren't set, instead of Caddy's own less obvious error from an empty{$DOMAIN}site address.
Caddy's on-disk certificate storage layout is a stable, documented
convention: /data/caddy/certificates/acme-v02.api.letsencrypt.org-directory/<domain>/<domain>.crt
(and .key) for the production Let's Encrypt CA. nginx/docker-entrypoint.sh
reads this path directly (mounted read-only at /caddy-data) — there's
nothing else keeping the two containers in sync.
In the cloud (Cloudflare in front of nginx)
Production and staging (plan saas-launch, D19) keep nginx as the in-cluster router, since
the cookie to Authorization translation in nginx/auth.js and the per-path SSE timeouts
live there, and put the Cloudflare proxy (Free plan) in front for TLS, WAF and DDoS. Set
Cloudflare SSL mode to Full (strict).
Two ways to give the origin a certificate, and both are in use:
- Caddy + Let's Encrypt, the same mechanism as self-host: enable the compose
tlsprofile and setDOMAIN/ACME_EMAIL. HTTP-01 works through the Cloudflare proxy because Cloudflare forwards the challenge path to the origin — this is whattest.reelbolt.airuns. See "Renewals" below for the one non-obvious requirement. - A Cloudflare Origin CA certificate for
<domain>and*.<domain>, pointed at withTLS_CERT_FILE/TLS_KEY_FILE(mounted files; this overrides the Caddy and self-signed choice indocker-entrypoint.sh). Pick this when Cloudflare's SSL mode is Full (strict) and you would rather Cloudflare itself vouched for the origin, or when you need a wildcard. A wildcard matters because therunner.hostname below is a second name on the same listener; a single-name Let's Encrypt certificate does not cover it.
nginx behaviour is controlled by environment variables, all defaulting to the single-host values so local compose behaves as before:
| Variable | Default | Meaning |
|---|---|---|
UPSTREAM_GO_API, UPSTREAM_INFERENCE_API, UPSTREAM_WORKFLOW_ENGINE, UPSTREAM_WEB, UPSTREAM_SITE, UPSTREAM_DOCS, UPSTREAM_CADDY | the compose service URLs (http://workflow-engine:8080, ...) | Upstream base URLs, substituted into /etc/nginx/conf.d/common.inc at start. The cloud control VM points UPSTREAM_WORKFLOW_ENGINE at http://$ENGINE_PRIVATE_IP:8080. |
REAL_IP_FROM_CLOUDFLARE | false | true installs nginx/cloudflare-realip.conf: set_real_ip_from for Cloudflare's published ranges and real_ip_header CF-Connecting-IP. Needed for per-IP rate limits (C5) and logs behind the proxy. Only connections from those ranges are trusted, so a direct hit on the origin cannot spoof an address. Refresh the ranges from cloudflare.com/ips-v4 and /ips-v6 when they change. Leave off locally. |
CLIENT_MAX_BODY_SIZE | 100M | Request body limit for the main server blocks. |
UPSTREAM_RUNNER_GATEWAY, RUNNER_HOSTNAME | empty, runner.$DOMAIN | See below. |
TLS_CERT_FILE, TLS_KEY_FILE | empty | Operator-supplied certificate, e.g. the Origin CA cert. |
runner.
docs/runner-protocol.md has runners dial wss://runner.<domain>/v1/connect. When
UPSTREAM_RUNNER_GATEWAY is set, the entrypoint renders nginx/runner-server.conf.template:
a separate server block for RUNNER_HOSTNAME that proxies only /v1/connect to the
gateway with WebSocket upgrade headers, without the cookie translation (runners
authenticate with a signed Ed25519 hello inside the protocol) and with any client
Authorization header cleared, proxy_buffering off, and 120 s read/send timeouts. The
protocol heartbeats every 15 s and declares a runner dead after 45 s, so the proxy never
cuts a healthy idle socket; Cloudflare's own WebSocket idle limit (100 s) is also above 45 s.
Everything else on that hostname is a 404. The gateway does not exist yet (decision D5a), so
the variable is unset in the cloud override and no block is rendered (an empty file is
included). TODO(runner-gateway WP): add the runner-gateway service and set
RUNNER_GATEWAY_UPSTREAM in the control VM's .env, create the proxied runner DNS record,
and run the one-hour WebSocket soak from F15's acceptance test. If UPSTREAM_RUNNER_GATEWAY
is set with neither RUNNER_HOSTNAME nor DOMAIN, nginx refuses to start.
Renewals
This deployment obtains its certificate with Caddy over ACME HTTP-01, which works through
the Cloudflare proxy: Cloudflare forwards /.well-known/acme-challenge/ to the origin, where
nginx proxies it to the caddy container. Two things are worth knowing before changing any of it.
Cloudflare SSL mode must be Full or Full (strict) — never Flexible. In the Full modes
Cloudflare reaches the origin on :443; in Flexible it fetches :80. With
FORCE_HTTPS=true that is an infinite redirect loop, because nginx answers every :80
request with a 301 to https:// that Cloudflare then re-requests over plain HTTP. Check which
mode is in use by sampling the established sockets on the origin while driving traffic through
the edge:
ss -Htn state established | awk '{print $4}' | grep -oE ':(80|443)$' | sort | uniq -c
:443 means Full or Full (strict), and FORCE_HTTPS is then safe. Full (strict) is the
right target: it makes Cloudflare validate the origin certificate, so an attacker cannot
impersonate the origin to the edge.
The Host header on the challenge location is load-bearing. Caddy matches an incoming
challenge to the site it manages by hostname, so it only answers while it sees
Host: <DOMAIN>. nginx's default for proxy_pass is Host: $proxy_host — here the upstream
name, the bare word caddy, which Caddy does not manage. The request then falls through to
Caddy's own HTTP→HTTPS redirect (a 308 to https://caddy/.well-known/...) and Let's Encrypt
refuses that URL outright: invalid host: must end in IANA registered TLD.
This failure is invisible until renewal. A certificate obtained before FORCE_HTTPS was
switched on keeps working, so the site looks healthy for ~60 days and then the certificate
simply expires. nginx/http-server-dev.conf has always forwarded Host;
nginx/http-server-redirect.conf did not, and only the redirect block is used once
FORCE_HTTPS=true. It now forwards Host and the forwarding headers, matching the dev block.
Prove renewal rather than assume it. A cached ACME authorization will happily hide the fault, so clear the account as well as the certificate, and keep a copy first:
docker exec reelbolt-caddy-1 sh -c 'mv /data/caddy/certificates /data/caddy/certificates.bak; rm -rf /data/caddy/acme'
docker restart reelbolt-caddy-1
docker logs reelbolt-caddy-1 2>&1 | grep -E 'served key authentication|authz_status|certificate obtained|challenge failed'
served key authentication followed by certificate obtained successfully is a pass. Restore
certificates.bak if it is not. nginx reads the certificate only at start or on reload, so
after a re-issue run docker exec reelbolt-nginx-1 nginx -s reload or recreate the container —
and note that if nginx starts while no Caddy certificate exists yet it silently falls back to
the self-signed one, and keeps serving that until it is restarted.
Upload size
Cloudflare Free rejects request bodies over 100 MB, and nginx's own limit is 100M, while F15's
acceptance asks for a 500 MB upload. Options considered:
- Presigned direct-to-R2 browser uploads (the browser PUTs to R2, nginx and Cloudflare's proxy never see the bytes). Cleanest and also removes upload load from the control VM, but it needs an Inference API endpoint that issues presigned URLs, an R2 CORS policy, and a client change. Recommended, as a future work package.
- An unproxied (DNS-only)
upload.<domain>hostname straight to nginx. Cheap in nginx, but it exposes the origin IP, bypasses WAF/DDoS for that name, and needs the auth cookie to be valid on a second host (Domain=onreelbolt_token, which does not exist today) plus CORS from the app host. - Pay for a Cloudflare plan with a larger cap.
Decision for the beta: keep the 100 MB cap (documented limit; larger sources wait for option 1).
Only the config side ships now: CLIENT_MAX_BODY_SIZE lets an operator raise the nginx limit
if the proxy cap is later lifted or an unproxied host is added. No upload. server block is
implemented, because without the cookie-domain and CORS work it would not authenticate.
SSH
Deploy SSH from GitHub-hosted runners and its firewall and sshd hardening are described in
docs/ci.md ("SSH reachability for deploys").