Skip to main content

Two-box production deployment

Operational reference for the live two-box ReelBolt deployment: what runs on each machine, why the split is where it is, how to bring it up, how to operate it at 3am, and what is still broken.

Every value here was read off the running boxes rather than from a plan. Where the deployment and an older description disagree, the box wins, and the disagreement is called out inline.

This is not the deployment path the repository ships. infra/compose/ (docker-compose.cloud.yml, host.sh, env.*.example) is the designed cloud path: SHA-pinned cosign-signed GHCR images, a reelbolt.service systemd unit, and a .deploy.env / .deploy.prev rollback. This deployment uses none of that. It builds every image from source on the host and has no rollback mechanism at all. See What production would still require.

Verified 2026-10-08 against commit b1900986 on both hosts.

The two boxes at a glance​

Control planeMedia plane
Hostnamevm-reelbolt-control-prod-weuvm-reelbolt-media-prod-weu
Public IP2.29.46.2265.109.128.5
Private IP (Hetzner)10.0.0.210.0.0.3
SSHssh -i ~/.ssh/reelbolt_deploy [email protected]ssh -i ~/.ssh/reelbolt_deploy [email protected]
Architecturex86_64 (amd64)aarch64 (arm64)
CPU4 cores8 cores
RAM7751 MiB (7.6 GiB)15578 MiB (15.2 GiB)
Disk75 G, 11 G used (14%)150 G, 8.6 G used (6%)
OS / kernelUbuntu 24.04.1 LTS, kernel 6.8.0-52-genericUbuntu 24.04.1 LTS, kernel 6.8.0-52-generic
Docker / Compose27.5.1 / v2.32.427.5.1 / v2.32.4
Code/opt/reelbolt, branch saas-launch/opt/reelbolt, branch saas-launch
Running containers113
Local images12 (all amd64)5 (all arm64)
.env127 keys, mode 600 root:root56 keys, mode 600 root:root

Both hosts are checked out at b1900986. They are two commits behind the local saas-launch branch (396bee85): the boxes have not been updated since the legal-placeholder commit.

The two machines talk only over the Hetzner private network (10.0.0.0/16). The media box has no inbound public port except sshd on 22.

Why the split is where it is​

The split is not a load-balancing decision. It is a blast-radius decision, and three constraints force it.

Data services and the front door stay on x86​

Postgres, Garage (S3), RabbitMQ and Qdrant hold every user's data. The render stack runs untrusted, model-authored code: an agent writes Remotion and ffmpeg workloads that execute in a container. Keeping those two things on different machines means a sandbox escape lands on a host that holds no database.

The front door (nginx) lives with the data because that is where the TLS terminator, the ACME client and the single public 80/443 already are, and because every backend it proxies to other than the engine is on the same host.

The render stack runs on the ARM box because it builds there natively​

The media box is aarch64. Every image it runs is built on that box from source; nothing is pulled. This is not a preference, it is the only way those images exist at all. See the architecture constraint.

Only the sandbox-executor holds the Docker socket​

sandbox-executor mounts /var/run/docker.sock because it starts the sandbox containers that run untrusted render workloads. That mount is host root. It is the single most dangerous mount in the deployment, and it exists on exactly one container on one host, a host that holds no database.

Verified: on the control plane no container mounts the Docker socket at all (the autoheal service, which would, is pinned behind a never-activated profile). On the media plane exactly one does:

MOUNTS-DOCKER-SOCK: /reelbolt-sandbox-executor-1
/var/run/docker.sock -> /var/run/docker.sock
/var/lib/reelbolt/sandboxes -> /var/lib/reelbolt/sandboxes

What runs where​

Control plane, 10.0.0.2 (11 services)​

cd /opt/reelbolt
docker compose -f docker-compose.yml -f docker-compose.control.yml up -d
ServiceImageArchPublished on
postgrespostgres:16-alpineamd64127.0.0.1:5432, 10.0.0.2:5432
garagedxflrs/garage:v2.4.1amd64127.0.0.1:9000, 10.0.0.2:9000
rabbitmqrabbitmq:3-management-alpineamd6410.0.0.2:5672, mgmt 127.0.0.1:15672
qdrantqdrant/qdrant:latestamd6410.0.0.2:6333, 10.0.0.2:6334
inferencereelbolt-inference:latest (built)amd6410.0.0.2:8080
go-apireelbolt-go-api:latest (built)amd64none (in-network)
webreelbolt-web:latest (built)amd64none (in-network)
sitereelbolt-site:latest (built)amd64none (in-network)
docsreelbolt-docs:latest (built)amd64none (in-network)
mcpreelbolt-mcp:latest (built)amd64none (in-network)
nginxreelbolt-nginx:latest (built)amd640.0.0.0:80, 0.0.0.0:443

Also present: alpine:3 (amd64), pulled by the build stages.

Note the container naming. postgres sets an explicit container_name, so it is reelbolt-postgres; every other service is Compose-generated as reelbolt-<service>-1. Use docker compose ... exec <service> rather than hardcoding container names.

Nothing is published on the public IP except nginx. The data services bind CONTROL_PRIVATE_IP (10.0.0.2) and loopback. The override uses Compose's :? form, so an unset value fails loudly instead of silently binding 0.0.0.0.

Media plane, 10.0.0.3 (4 defined, 3 running)​

cd /opt/reelbolt
docker compose -f docker-compose.yml -f docker-compose.media.yml up -d
ServiceImageArchStatus
workflow-enginereelbolt-workflow-engine:latest (built)arm64Up, healthy. Publishes 10.0.0.3:8080
sandbox-executorreelbolt-sandbox-executor:local (built)arm64Up, healthy. Mounts the Docker socket
runner-gatewayreelbolt-runner-gateway:local (built)arm64Up, healthy. In-network 8090
sandbox-runtimereelbolt-sandbox-runtime:local (built)arm64Exited (0), see below
SERVICE STATUS
runner-gateway Up 8 minutes (healthy)
sandbox-executor Up 9 minutes (healthy)
sandbox-runtime Exited (0) 4 minutes ago
workflow-engine Up 8 minutes (healthy)

sandbox-runtime is not a failure. It is a deliberate build-and-exit service: it builds the image that untrusted render code runs inside, then stops. Its own compose comment says so: "Build-and-exit: produces the image untrusted, model-authored render code runs in." An Exited (0) for this service is the healthy state, and docker compose ps omitting it from the running list is expected. Do not "fix" it.

Disagreement with the brief. The brief describes four running services on the media plane. Four are defined; only three run. sandbox-runtime is a one-shot build stage.

The image it produces is what the executor launches:

SANDBOX_IMAGE=reelbolt-sandbox-runtime:local

The profiling mechanism that keeps the hosts honest​

Neither host edits docker-compose.yml. Both add a pure override layer, and each pins the other plane's services behind a never-activated profile:

  • Control pins workflow-engine, sandbox-runtime, sandbox-executor and runner-gateway behind profile media-plane.
  • Media pins all eleven control services behind profile media-disabled.
  • Both pin whisper, embeddings and autoheal off. Those have no profile in the base file, so without pinning them they would start.

The media file moves its control-plane services behind a profile, quoting its own comment, "rather than deleted, so a bare docker compose up on this host can never start a second, split-brain Postgres/RabbitMQ/Garage or a second API." That is the guardrail: always pass both -f flags.

Bring-up from bare metal​

The control plane has a one-shot script; the media plane does not.

1. Both hosts: prerequisites​

  • Ubuntu 24.04, Docker 27.5.1, Compose v2.32.4.
  • The deployment key ~/.ssh/reelbolt_deploy authorized in /root/.ssh/authorized_keys.
  • The Hetzner private network up, so 10.0.0.2 and 10.0.0.3 are mutually reachable.
  • /opt/reelbolt on branch saas-launch.

2. Control plane​

ssh -i ~/.ssh/reelbolt_deploy [email protected]
cd /opt/reelbolt
# .env must exist first: mode 600, see the .env contract below
./bringup.sh

bringup.sh (15 lines, idempotent, re-runnable) is:

COMPOSE="docker compose -f docker-compose.yml -f docker-compose.control.yml"
$COMPOSE build
$COMPOSE up -d
$COMPOSE ps

It is untracked in git: it exists only on the box. Build the control plane first. The Inference API owns the schema and migrates on startup, and the engine on the other host migrates after it.

3. Media plane​

There is no bringup.sh here. Run the two commands by hand:

ssh -i ~/.ssh/reelbolt_deploy [email protected]
cd /opt/reelbolt
docker compose -f docker-compose.yml -f docker-compose.media.yml build
docker compose -f docker-compose.yml -f docker-compose.media.yml up -d

Build here, before up. The build is what produces the arm64 images; there is nothing to pull.

4. The media box has no GitHub credential, so git pull fails there​

Verified on the box:

cd /opt/reelbolt && git fetch --dry-run
# fatal: could not read Username for 'https://github.com': No such device or address

The remote is HTTPS (https://github.com/vecchiotom/ReelBolt.git), there is no credential helper configured, and /root/.ssh contains only authorized_keys and known_hosts, with no deploy key. The code reached that box out of band.

Workaround (what was actually done): transfer the tree from the control plane. From the control box:

cd /opt/reelbolt && git bundle create /tmp/reelbolt.bundle saas-launch
scp -i ~/.ssh/reelbolt_deploy /tmp/reelbolt.bundle [email protected]:/tmp/

Then on the media box:

cd /opt/reelbolt
git fetch /tmp/reelbolt.bundle saas-launch:refs/remotes/bundle/saas-launch
git merge refs/remotes/bundle/saas-launch

The real fix is a read-only deploy key. A GitHub deploy key scoped to this repository, installed in the media box's /root/.ssh/, plus an SSH remote ([email protected]:vecchiotom/ReelBolt.git). See What still needs the owner. Until then the media box cannot be deployed to independently, and any update must be pushed from the control plane.

Disagreement with the brief. The brief says both boxes run code at /opt/reelbolt on branch saas-launch without qualification. True, but the media checkout is not deployable on its own, and neither checkout is clean: bringup.sh and docker-compose.control.yml are untracked on the control box, and docker-compose.media.yml is untracked on the media box. Those override files exist nowhere in git. Back them up separately: git clean -fdx would destroy the deployment.

5. What starts the stack after a reboot​

There is no systemd unit. Verified on both hosts: no reelbolt.service. What brings the stack back after a reboot is Docker's own restart policy — every container is unless-stopped, and docker.service is enabled, so the daemon restarts them all by itself.

The consequence for the 3am reader: containers come back by themselves; new code does not. A reboot restores the running stack, not a deployed one.

Since 2026-10-09 there is a release mechanism and a rollback target: infra/scripts/release.sh records the live commit in .deploy.sha and the outgoing one in .deploy.prev. See "Releases and rollback" below. (This replaces bringup.sh, which ran compose build && compose up -d and left no record of what had been running, so a bad release could not be undone. bringup.sh is no longer tracked and should not be used.)

Releases and rollback​

A release is a git commit. infra/scripts/release.sh runs on the box and is the only thing that should ever start this stack by hand.

# On either box, from /opt/reelbolt:
REELBOLT_ROLE=control bash infra/scripts/release.sh deploy # release the current checkout
REELBOLT_ROLE=control bash infra/scripts/release.sh deploy <sha> # fetch and release a specific commit
REELBOLT_ROLE=control bash infra/scripts/release.sh rollback # back to .deploy.prev
REELBOLT_ROLE=control bash infra/scripts/release.sh status
REELBOLT_ROLE=control bash infra/scripts/release.sh prune 5 # keep the 5 newest releases

REELBOLT_ROLE is control or media and selects both the override layer (docker-compose.control.yml / docker-compose.media.yml) and the health checks: control probes https://127.0.0.1/{health,api/v1/health,api/v1/workflow-engine/health} through nginx, so a green release means the front door and everything behind it answer; media waits on the engine's own container healthcheck.

How a release is recorded​

.deploy.sha the release that is live now
.deploy.prev the release that was live before it

Deploying tags every locally-built image as <repository>:<sha>, so the previous release's images survive the next one. rollback re-points the reference compose starts each container from at the previous tag and restarts — no rebuild, so it takes as long as a restart. A failed health check performs exactly that rollback automatically and then exits non-zero.

Two things about image names, because both have already caused a failed release:

  • The reference is not always <project>-<service>. The base compose pins reelbolt-sandbox-runtime:local, reelbolt-sandbox-executor:local and reelbolt-runner-gateway:local; the rest are left for compose to name, giving reelbolt-<service>:latest. release.sh reads the composed config rather than assuming, and keeps the repository and the tag separate — using the tagged reference as the base for the release tag produces reelbolt-workflow-engine:latest:<sha>, which Docker rejects outright.
  • Schema is not rolled back. Migrations only move forward, so a rollback restores code, not the database. Write migrations expand/contract.

Both boxes hold a read-only GitHub deploy key at /root/.ssh/reelbolt_github (registered on the repo as "two-box deploy (read-only)"), with GitHub's host key pinned in /root/.ssh/known_hosts, so deploy <sha> can fetch a commit that is not on the box yet.

Retiring Garage​

Object storage moved to Cloudflare R2 (2026-10-09) and Garage is pinned to the cloud-disabled profile so compose up -d cannot resurrect it. The reelbolt_garagedata/reelbolt_garagemeta volumes are kept deliberately, so the migration stays reversible and the old objects remain readable. go-api's depends_on had to be overridden to drop Garage: a depends_on naming a service that no active profile defines invalidates the entire compose project, with the unhelpful message service "go-api" depends on undefined service "garage".

The config keys are still spelled MinIO__* / MINIO_*; that is historical. What matters for R2 is MINIO_REGION=auto (Garage enforced garage) and path-style addressing.

Certificate renewal​

Caddy obtains and renews the certificate; nginx only reads it. The one non-obvious requirement — the Host header on the ACME challenge location, without which renewal fails silently for ~60 days — is in tls.md under "Renewals". Read that before touching nginx/http-server-redirect.conf.

The .env contract​

Both hosts read /opt/reelbolt/.env, mode 600 root:root. These files hold every credential the deployment has. Never print their contents, never paste them into a ticket, never commit them. Every value in this section is referenced by variable name only.

  • Control: 127 keys, 35 # PASTE: markers.
  • Media: 56 keys, 11 # PASTE: markers.

Variables that must be byte-identical across the two hosts​

This is the part that bites. The two services share a Postgres database, a RabbitMQ broker and an object store, and the Go API issues JWTs that the engine validates. Each value below was compared across the two boxes by hashing it and comparing digests; every one is confirmed identical. Changing any of them on one host only breaks the deployment.

VariableWhy it must match
JWT_SIGNING_KEYThe Go API signs; the Inference API and the engine verify. A mismatch invalidates every session token and every engine-to-API call.
POSTGRES_USER, POSTGRES_PASSWORD, POSTGRES_DBOne database, two hosts.
RABBITMQ_USER, RABBITMQ_PASSWORDThe engine consumes from the control plane's broker.
MINIO_ACCESS_KEY, MINIO_SECRET_KEY, MINIO_BUCKETOne Garage bucket, written by both hosts.
SANDBOX_API_TOKENThe engine authenticates to the sandbox executor with it.
RUNNER_INTERNAL_TOKENEngine to runner-gateway internal auth.
RUNNER_JOB_TOKEN_KEYSigns runner job tokens; minted on one side, verified on the other.
CONTROL_PRIVATE_IP (10.0.0.2)Each host needs the other's address; both hold both.
ENGINE_PRIVATE_IP (10.0.0.3)As above.
TENANCY_MODECurrently Cloud on both. A split here makes the two services disagree about tenancy.

Present on the control plane but not on the media plane​

  • INTERNAL_API_TOKEN: set on control, absent on media. The media .env does not define it at all. Nothing on the media plane currently reads it, so this is not an active fault, but it is an asymmetry to know about before wiring anything new that expects it.
  • The entire front-door and third-party block: DOMAIN, PUBLIC_BASE_URL, SIGNUP_ENABLED, OAUTH_*, SMTP_*, PADDLE_*, TURNSTILE_*, PLATFORM_ASR_*, PLATFORM_EMBEDDING_* and the rest.

Present on the media plane but not on the control plane​

  • DATABASE_CONNECTION_STRING: the .NET connection string, pointing at Host=10.0.0.2;Port=5432.
  • DATABASE_URL: points at 10.0.0.2.
  • MINIO_ENDPOINT, MINIO_REGION, MINIO_FORCE_PATH_STYLE: the control plane omits these and relies on the base compose defaults; the media override requires them.
  • DATA_PROTECTION_STORE, AWS_REGION: see the defect below. Both are currently inert.
  • The media-only render tuning keys: VIDEO_FFMPEG_THREADS, OTEL_*.

Both hosts: DATABASE_URL and DATABASE_CONNECTION_STRING embed the Postgres password in plaintext​

This is a real finding, stated without the value. Every .env also carries POSTGRES_PASSWORD separately, so the same credential appears twice in the control file and three times in the media file.

Rotating the Postgres password therefore means editing five places across two hosts, and missing one produces a failure that looks like a network problem. Change POSTGRES_PASSWORD, DATABASE_URL and DATABASE_CONNECTION_STRING together on the media box, and POSTGRES_PASSWORD and DATABASE_URL on the control box. For reference, the connection strings are shaped like:

DATABASE_URL=postgres://<user>:<REDACTED>@10.0.0.2:5432/reelforge?sslmode=disable
DATABASE_CONNECTION_STRING=Host=10.0.0.2;Port=5432;Database=reelforge;Username=<user>;Password=<REDACTED>

The PASTE markers​

Both files ship from .env.example with every third-party credential blank and marked # PASTE:.

  • Control: 35 markers. Azure OpenAI (endpoint and key), Anthropic (Console key and subscription token), Google AI Studio and Google Cloud keys, DeepSeek key, the hosted ASR provider block (name, endpoint, model, key), the hosted embedding block (name, endpoint, model, key), SMTP (host, username, password, from), Turnstile (secret and site key), Google and GitHub OAuth (id and secret pairs), hosted Qdrant URL and key, ACME e-mail, GA4 measurement id, contact webhook URL, and the five Paddle values (API key, webhook secret, three top-up price ids).
  • Media: 11 markers. Five are copy-from-control markers (JWT_SIGNING_KEY, RabbitMQ, DATABASE_CONNECTION_STRING, DATABASE_URL, MinIO) and those are already filled. The remaining six are provider credentials (Anthropic, Google, DeepSeek, Azure OpenAI) plus an OTEL authorization header, and are still blank.

A blank here is not always a fault. SIGNUP_ENABLED=false and BILLING_ENABLED=false mean the Paddle and Turnstile blocks are legitimately empty today. But PLATFORM_ASR_NAME and PLATFORM_EMBEDDING_NAME are both blank on the control plane, which means no transcription and no embedding provider is configured at all: the local whisper and embeddings containers are pinned off, so both capabilities currently have no source. Any workflow step needing ASR or semantic file search fails until an owner pastes one.

The architecture constraint: aarch64 vs x86_64​

The control plane is x86_64 (amd64). The media plane is aarch64 (arm64). They cannot share images. This is the constraint that shapes every deploy decision on the media box.

Verified: every image on the media box is arm64, and every image on the control plane is amd64.

for i in $(docker images --format '{{.Repository}}:{{.Tag}}' | sort -u); do
echo "$i -> $(docker image inspect --format '{{.Architecture}}/{{.Os}}' "$i")"
done
# media plane (aarch64 host)
alpine:3 arm64/linux
reelbolt-runner-gateway:local arm64/linux
reelbolt-sandbox-executor:local arm64/linux
reelbolt-sandbox-runtime:local arm64/linux
reelbolt-workflow-engine:latest arm64/linux
# control plane (x86_64 host)
alpine:3 amd64/linux
dxflrs/garage:v2.4.1 amd64/linux
postgres:16-alpine amd64/linux
qdrant/qdrant:latest amd64/linux
rabbitmq:3-management-alpine amd64/linux
reelbolt-docs:latest amd64/linux
reelbolt-go-api:latest amd64/linux
reelbolt-inference:latest amd64/linux
reelbolt-mcp:latest amd64/linux
reelbolt-nginx:latest amd64/linux
reelbolt-site:latest amd64/linux
reelbolt-web:latest amd64/linux

Every image the media box runs is built on the media box​

The media override states it directly: "The images are built HERE, on this ARM64 box, instead of pulled from GHCR. A plain docker build on aarch64 produces an arm64 image, which is what the control plane (x86_64) must never receive. Nothing amd64 is pulled."

So on the media plane the sequence is always build then up -d. There is no registry to fall back on and no IMAGE_TAG pin.

Verify the architecture of anything you rebuild​

docker image inspect --format '{{.Architecture}}' reelbolt-workflow-engine:latest
# must print: arm64

Check all four after any rebuild:

for i in reelbolt-workflow-engine:latest reelbolt-sandbox-executor:local reelbolt-sandbox-runtime:local reelbolt-runner-gateway:local; do
printf '%-42s %s\n' "$i" "$(docker image inspect --format '{{.Architecture}}' "$i")"
done

Why pulling an amd64 image there would fail, or silently emulate​

  • Fail: the kernel cannot execute amd64 ELF binaries. You get exec format error at container start, often deep inside a healthcheck loop, with the container restarting repeatedly.
  • Silently emulate: if binfmt_misc or QEMU handlers are registered on the host, Docker will happily run the amd64 image under emulation. Nothing errors. It is simply wrong: an arm64 host emulating amd64 runs an ffmpeg or Remotion render several times slower, at which point VIDEO_MAX_CONCURRENT_JOBS=2 oversubscribes 8 real cores and renders start timing out. This is the dangerous case, because the symptom is "the box got slow", not "the box is broken".

The sharp edge: SANDBOX_IMAGE=reelbolt-sandbox-runtime:local is what sandbox-executor launches for every untrusted render. If a rebuild ever produces an amd64 reelbolt-sandbox-runtime:local, every sandbox container is wrong-arch. Sandbox containers are created on demand, so this never shows up in docker compose ps. It shows up as failed renders.

Day-2 operations​

All commands run from /opt/reelbolt with the correct pair of -f flags. Define them once:

# control plane
alias rbc='docker compose -f docker-compose.yml -f docker-compose.control.yml'
# media plane
alias rbm='docker compose -f docker-compose.yml -f docker-compose.media.yml'

Restart one service​

# control plane
cd /opt/reelbolt
docker compose -f docker-compose.yml -f docker-compose.control.yml restart go-api

# media plane
cd /opt/reelbolt
docker compose -f docker-compose.yml -f docker-compose.media.yml restart workflow-engine

On the control plane a restart can cascade, and this is expected. nginx declares depends_on with restart: true for site, web, go-api and inference. Restarting any of those also restarts nginx, a brief blip on the public front door. Verified:

$ docker compose -f docker-compose.yml -f docker-compose.control.yml restart go-api
Container reelbolt-go-api-1 Restarting
Container reelbolt-go-api-1 Started
Container reelbolt-nginx-1 Restarting <-- cascade
Container reelbolt-nginx-1 Started

The media plane does not cascade: restarting workflow-engine leaves runner-gateway and sandbox-executor untouched.

restart does not rebuild, which is the common case. To pick up new code, use up -d --build <service>.

Read logs​

# follow one service on the control plane
docker compose -f docker-compose.yml -f docker-compose.control.yml logs -f --tail=200 inference

# media engine
docker compose -f docker-compose.yml -f docker-compose.media.yml logs -f --tail=200 workflow-engine

# since a timestamp
docker compose -f docker-compose.yml -f docker-compose.control.yml logs --since 30m go-api

By container name when you already know it. Remember reelbolt-postgres is the one exception to the -1 suffix rule:

docker logs --tail=200 reelbolt-postgres
docker logs --tail=200 reelbolt-workflow-engine-1

The engine is the busiest log in the system. Start there when a workflow misbehaves.

Check health​

The authoritative source is the container healthchecks themselves. Read what the box actually runs rather than guessing a path:

for c in $(docker ps -q); do
n=$(docker inspect -f '{{.Name}}' $c | sed 's|^/||')
echo "$n :: $(docker inspect -f '{{.State.Health.Status}}' $c 2>/dev/null || echo no-healthcheck)"
done

Real healthcheck endpoints on this deployment:

ServiceHealthcheck
nginxhttps://127.0.0.1/health and https://127.0.0.1/api/v1/health
go-apihttp://localhost:8080/health
inferencehttp://localhost:8080/api/v1/health
webhttp://127.0.0.1:3000/app/login
sitehttp://127.0.0.1:3000/
docshttp://127.0.0.1/healthz
mcphttp://localhost:3002/health
postgrespg_isready -U postgres -d reelforge
garage/garage status
rabbitmqrabbitmq-diagnostics -q check_running
workflow-engine (media)http://localhost:8080/api/v1/workflow-engine/health
sandbox-executor (media)http://localhost:8080/health
runner-gateway (media)http://localhost:8090/healthz

qdrant has no healthcheck. It shows as Up with no health state, which is normal.

The engine's health endpoint is at /api/v1/workflow-engine/health, not /health. The bare paths 404. Verified through the private network from the control plane:

curl -s http://10.0.0.3:8080/api/v1/workflow-engine/health
# {"status":"healthy","service":"workflow-engine","timestamp":"..."}

Front-door check. Note -k: the certificate is self-signed.

curl -sk -o /dev/null -w '%{http_code}' https://2.29.46.22/ ; echo
curl -sk -o /dev/null -w '%{http_code}' https://2.29.46.22/api/v1/health ; echo
curl -sk -o /dev/null -w '%{http_code}' https://2.29.46.22/api/v1/workflow-engine/health ; echo

Back up Postgres​

Postgres holds all application state: users, organizations, projects, workflows, executions, metering, billing and the schema itself. This is the only backup that really matters.

# on the control plane
docker exec reelbolt-postgres pg_dump -U postgres -d reelforge -Fc > /root/reelforge-$(date -u +%Y%m%dT%H%M%SZ).dump

Verified working: pg_dump (PostgreSQL) 16.15, database size 10207 kB, a plain-SQL dump of about 302 KB. Small, but the billing and metering ledgers cannot be reconstructed from anywhere else.

Restore:

docker exec -i reelbolt-postgres pg_restore -U postgres -d reelforge --clean --if-exists < /root/reelforge-YYYYMMDDTHHMMSSZ.dump

pg_dump alone is not sufficient here. See the Data Protection defect below: restoring the database without the matching key ring leaves every encrypted provider key permanently unreadable.

Back up the Garage bucket​

Garage is the S3-compatible object store. The bucket is reelforge.

docker exec reelbolt-garage-1 /garage bucket list
docker exec reelbolt-garage-1 /garage key list
docker exec reelbolt-garage-1 /garage status

Verified: one node 76537915d73d7d9f, capacity 74.8 GiB with 61.7 GiB available, bucket reelforge, one key. Object data is currently tiny, about 8 KB of volume data, because the stack holds no production media yet.

The Garage container has no shell. sh is not in the image, so docker exec ... sh -c fails with exec: "sh": executable file not found. Use the /garage binary directly, as above.

Garage provides no backup subcommand. Back it up at the filesystem level, with the container stopped so nothing is mid-write:

# on the control plane
docker compose -f docker-compose.yml -f docker-compose.control.yml stop garage
tar czf /root/garage-$(date -u +%Y%m%dT%H%M%SZ).tar.gz -C /var/lib/docker/volumes reelbolt_garagedata reelbolt_garagemeta
docker compose -f docker-compose.yml -f docker-compose.control.yml start garage

The two volumes are confirmed as:

  • /var/lib/docker/volumes/reelbolt_garagedata/_data mounted at container /data (objects)
  • /var/lib/docker/volumes/reelbolt_garagemeta/_data mounted at container /meta (metadata)

You need both. Restoring /data without /meta gives you blocks with no index, and garage repair blocks is an involved recovery, not a quick fix. Garage also stores the cluster rpc_secret in /opt/reelbolt/garage/garage.toml. Back that up too, or the restored node has a different identity.

This is the single most important check when something "cannot reach the database". Run all of it from the control plane.

1. RabbitMQ: which hosts are connected to the broker. The media engine's connections appear with peer_host 10.0.0.3:

docker exec reelbolt-rabbitmq-1 rabbitmqctl list_connections name peer_host

Observed:

name peer_host
172.18.0.8:44112 -> 172.18.0.5:5672 172.18.0.8 <- in-network (control)
172.18.0.9:59470 -> 172.18.0.5:5672 172.18.0.9 <- in-network (control)
10.0.0.3:59256 -> 172.18.0.5:5672 10.0.0.3 <- MEDIA ENGINE
10.0.0.3:42154 -> 172.18.0.5:5672 10.0.0.3 <- MEDIA ENGINE

Two connections from 10.0.0.3 is the proof the cross-box link is up. If they are missing, the engine is down, the private network is broken, or the credentials diverge.

Queue state. The engine is the consumer:

docker exec reelbolt-rabbitmq-1 rabbitmqctl list_queues name messages consumers

Observed: ProjectFileIndexing, workflow-execution, workflow-stop-requests and go-api-workflow-events, each with a consumer and 0 messages backed up.

2. Postgres: which hosts hold connections. The media engine's pool appears as 10.0.0.3:

docker exec reelbolt-postgres psql -U postgres -d reelforge -c "select client_addr, usename, datname, state, count(*) from pg_stat_activity where backend_type = 'client backend' group by 1,2,3,4 order by 1;"

Observed:

client_addr | usename | datname | state | count
-------------+----------+-----------+--------+-------
10.0.0.3 | postgres | reelforge | idle | 5 <- MEDIA ENGINE POOL
172.18.0.8 | postgres | reelforge | idle | 1 <- inference (in-network)
172.18.0.9 | postgres | reelforge | idle | 1 <- go-api (in-network)
| postgres | reelforge | active | 1 <- this query

Five idle connections from 10.0.0.3 is the proof the engine reaches the database. The filter must be backend_type = 'client backend'; with an underscore the query errors out.

3. Raw TCP reachability. If the two checks above fail, this says whether it is the network or the application:

# from the MEDIA box, toward the control plane
for hp in 10.0.0.2:5432 10.0.0.2:5672 10.0.0.2:9000 10.0.0.2:6333 10.0.0.2:8080; do
h=${hp%:*}; port=${hp#*:}
if timeout 5 bash -c "cat < /dev/null > /dev/tcp/$h/$port" 2>/dev/null; then echo "OPEN $hp"; else echo "CLOSED $hp"; fi
done

All five are currently OPEN: Postgres, RabbitMQ, Garage, Qdrant REST and the Inference API.

4. Confirm the media box still has no inbound public port:

ss -tln | grep -v '127.0.0.1' | grep -v '10.0.0.3'

Expected: only sshd.

LISTEN 0 4096 *:22 *:*

Known defects and their current state​

Re-check these before acting on them. At the time of writing two of the three were still live, and another agent may have fixed one since. Every claim below is what was observed on the boxes on 2026-10-08, not what was planned.

1. The Data Protection key ring is per-host: STILL BROKEN​

Status: open. Currently latent, because no encrypted keys are stored yet, but guaranteed to bite the first time an admin types an API key into the UI.

The Inference API encrypts provider API keys at rest with ASP.NET Core Data Protection (ISecretProtector). The engine decrypts what the API wrote, so both services must share one key ring. They do not.

Each host has its own dpkeys volume, and the two volumes contain different keys:

# control plane
reelbolt_dpkeys/key-7aa7e137-1680-4bea-b7c0-f35fc4270a76.xml

# media plane
reelbolt_dpkeys/key-f9792ecb-2b94-482e-b7fe-823c58e5e5bf.xml

Two different key-ring files on two hosts. Neither can read the other's ciphertext.

Why it is latent right now: the inference_providers table is empty, 0 rows on the deployment. Nothing has been encrypted yet. Keys whose value is the literal env: sentinel are unaffected, because they are never persisted.

The trap: the failure mode is not a startup error. It appears at first use: a provider row saved on the control plane, then a workflow step on the media engine that tries to decrypt it and fails. Nothing in docker compose ps will look wrong.

The evidence that the code is wired correctly but the configuration is not: the base docker-compose.yml sets only DataProtection__KeysPath: /keys (inference at line 331, engine at line 443). It sets no DataProtection__Store. The repository's own cloud override infra/compose/docker-compose.cloud.yml does set DataProtection__Store: ${DATA_PROTECTION_STORE:-Database}, with the comment: "a per-host /keys directory would leave the other VM unable to read keys this one wrote." The designed cloud path solves this; this deployment does not use it.

A red herring to be aware of: the media .env contains DATA_PROTECTION_STORE=FileSystem. That variable is inert; nothing reads it. DATA_PROTECTION_STORE appears nowhere in docker-compose.yml, and the engine container's environment contains only DataProtection__KeysPath=/keys. Editing that line changes nothing. AWS_REGION in the media .env is likewise a leftover.

What the fix requires: set DataProtection__Store=Database on both the Inference API and the engine, plus a shared wrapper key: either DATA_PROTECTION_KMS_LOCAL_KEY (a base64 32-byte value, identical on both hosts) or AWS KMS via DATA_PROTECTION_KMS_KEY_ID plus credentials. Then restart both services. See docs/secrets.md and infra/compose/docker-compose.cloud.yml for the exact keys. Do this before anyone stores a real provider key, because migrating an existing per-host ring is a manual re-wrap, not an automatic one.

2. Presigned URLs point at a private address: STILL SET, currently harmless​

Status: the misconfiguration is present; its impact is latent because the compute fabric is off.

MINIO_PUBLIC_ENDPOINT=http://10.0.0.2:9000 on both hosts: the private Hetzner address.

MinIO__PublicEndpoint is the host presigned URLs are signed against. SigV4 signs the Host header, so this value must be the address the consumer of the URL will actually connect to. docs/cloud-storage.md states the rule: "it must be the host a runner or browser really connects to." 10.0.0.2 is unroutable from anywhere outside the private network.

Why it is harmless today:

  1. COMPUTE_MODE=Local on both hosts. The only consumers of engine-minted presigned URLs are the runner-fabric paths (SandboxWorkspaceJournal, RemoteSandboxSession, OutputVerifier). With the fabric off they are never exercised.
  2. The browser never receives a presigned URL. Downloads are streamed through the Inference API (RangedObjectStreamer.StreamAsync over the API's own in-network S3 client), not redirected to the object store. So this defect cannot break a dashboard download today.

When it will bite: the moment COMPUTE_MODE=Fabric is enabled, or a desktop runner (docs/desktop-runner.md), a machine outside the private network, receives a manifest. Every presigned GET and PUT it is handed will name an address it cannot reach. The failure is a timeout or a connection refused inside the sandbox, not an auth error, which makes it look like a network problem.

A correctness note on the current value: because the runner and the sandbox containers all live on the media box, they can reach 10.0.0.2:9000, so a Fabric rollout would appear to work and then fail for exactly the first off-network consumer. If a real domain is provisioned, this value must change to the public HTTPS endpoint of the object store at the same time.

3. Tenancy mode: FIXED, verified Cloud on both hosts​

Status: correct. This is what the box shows, not what was planned.

TENANCY_MODE=Cloud is set in the .env of both hosts, is passed through to go-api by the control override, and is live in the running container:

docker exec reelbolt-go-api-1 printenv TENANCY_MODE
# Cloud

Corroborated by the data. The tenancy shape is the Cloud one, with no shared Default organization:

docker exec reelbolt-postgres psql -U postgres -d reelforge -c "select kind, count(*) from organizations group by kind;"
kind | count
----------+-------
Personal | 1
Team | 1
Platform | 1

A SelfHosted deployment would put users in a shared Default organization. There is none.

Why this variable must be passed explicitly. From the control override: the base compose file never sets TENANCY_MODE, so "a cloud deployment would silently run SelfHosted (users join a shared Default org)." If you ever rewrite the override, keep this line. Its absence is silent, and the symptom is a permissions model that is not the one you think you deployed.

Related and also correct: SIGNUP_ENABLED=false and NEXT_PUBLIC_SIGNUP_ENABLED=false.

What still needs the owner​

Four things cannot be done from inside the deployment, because each needs an external credential, a DNS record, or a decision that is not an engineer's to make.

1. A real domain: the stack runs on a bare IP with a self-signed certificate​

VariableValue
PUBLIC_BASE_URLhttp://2.29.46.22
NEXT_PUBLIC_SITE_URLhttp://2.29.46.22
DOMAINempty
ACME_EMAILempty
FORCE_HTTPSfalse
COOKIE_SECUREfalse

nginx serves a self-signed certificate:

subject=CN = localhost
issuer=CN = localhost
notBefore=Oct 8 22:55:23 2026 GMT
notAfter =Oct 8 22:55:23 2027 GMT

Consequences: browsers show a certificate warning (tls_verify_result=18, self-signed), there is no trusted TLS, and plain HTTP on port 80 serves the application directly with no redirect to HTTPS.

Disagreement with the brief. The brief says the stack runs on http://2.29.46.22 with a self-signed cert. Both halves are true simultaneously: port 80 and port 443 are both open and both return 200. That is worse than "HTTP only", because credentials can be sent in cleartext to a working endpoint while the HTTPS listener sits there ignored.

Needed: a domain and its A records, DOMAIN and ACME_EMAIL set in .env, and the documented cut-over in docs/tls.md, which activates the caddy container, already pinned behind the opt-in tls profile and left untouched for exactly this purpose.

2. Third-party credentials: 35 blank markers on control, 11 on media​

Every third-party integration is unconfigured. The full inventory is in the PASTE markers section. The ones that block a launch:

  • A transcription (ASR) provider and an embedding provider. PLATFORM_ASR_NAME and PLATFORM_EMBEDDING_NAME are blank, and the local whisper and embeddings containers are pinned off. The deployment currently has no speech-to-text and no semantic search at all.
  • A chat provider. Every chat key (ANTHROPIC_API_KEY, GOOGLE_API_KEY, DEEPSEEK_API_KEY, AZURE_OPENAI_API_KEY) is blank, so no agent can run.
  • SMTP. Password resets and email verification fall back to console logging, so no user can reset a password.
  • Paddle (key, webhook secret, three price ids), required before BILLING_ENABLED=true.
  • OAuth (Google, GitHub) and Turnstile: optional, currently all disabled.
  • ACME e-mail, required for TLS.

3. A read-only deploy key (GitHub)​

The media box has no GitHub credential and cannot pull. See the media box has no GitHub credential. Needed: a repository-scoped read-only deploy key, not a personal token and not a write key, installed on the media box, plus an SSH remote. Read-only is the right scope: the media box only ever needs to fetch.

4. TLS​

Covered above under the domain. docs/tls.md is the runbook; the caddy container and the nginx /.well-known/acme-challenge/ location are already wired and waiting.

Security posture and its rationale​

Why the media box has no inbound public port​

The media box runs untrusted, model-authored code. Its job is to execute code that an AI agent wrote, on behalf of a user, with no human review. Giving that host a public attack surface would be indefensible. It has one listening public port, sshd, and everything else is private:

LISTEN 0 4096 127.0.0.54:53 0.0.0.0:*
LISTEN 0 4096 127.0.0.53%lo:53 0.0.0.0:*
LISTEN 0 4096 10.0.0.3:8080 0.0.0.0:* <- engine, PRIVATE address only
LISTEN 0 4096 *:22 *:* <- the only public inbound

The engine publishes on 10.0.0.3:8080, never 0.0.0.0:8080. The control plane's nginx is the only caller (UPSTREAM_WORKFLOW_ENGINE=http://10.0.0.3:8080).

Why only the sandbox-executor mounts the Docker socket​

sandbox-executor must start containers, so it needs /var/run/docker.sock, which is host root. Anyone who escapes that container owns the host. The response is not to remove the mount, since the service cannot work without it, but to bound what the host holds:

  • The socket lives on the media box, which holds no database. The database is on the other host.
  • The mount is on exactly one container on that host. Verified on both boxes.
  • On the control plane nothing mounts the socket at all. The autoheal service, which would, is pinned behind the never-activated cloud-disabled profile, and the override says why: "With it off, NOTHING on this box mounts the Docker socket."
  • The control plane's web service has its 9229 Node inspector publish reset to empty, described in the override as "unauthenticated RCE for anyone who can reach the port."

This is the whole reason the two-box split exists. A sandbox escape on the media box reaches a host with no user data, no database and no object store.

Password authentication is still enabled on both hosts, deliberately​

Both hosts currently have:

permitrootlogin yes
pubkeyauthentication yes
passwordauthentication yes

Confirmed in /etc/ssh/sshd_config.d/50-cloud-init.conf on both.

This is a deliberate, temporary state. Key authentication is proven to work: auth.log shows 90 successful publickey logins on the control plane and 61 on the media plane. Password auth was left on as a fallback and the cut-over was never performed.

The weak point is not the key auth. It is that PermitRootLogin yes combined with password auth means the internet can attempt to brute-force root, and root's only constraint is PAM. There is no fail2ban configured on either host.

What production would still require​

In priority order:

  1. Turn off password authentication. Set PasswordAuthentication no in /etc/ssh/sshd_config.d/50-cloud-init.conf, or better PermitRootLogin prohibit-password, on both hosts, and validate the key-only path in a second session before closing the first. Key auth is already proven, so this is a five-minute change that removes an unbounded risk.
  2. Terminate real TLS. A domain plus the documented docs/tls.md cut-over. Also set COOKIE_SECURE=true and FORCE_HTTPS=true once HTTPS is the only listener. Today COOKIE_SECURE=false means session cookies are sent over plain HTTP.
  3. Configure backups. There are none. Neither host has a root crontab, and every systemd timer on both hosts is an OS default (dpkg-db-backup, logrotate, fstrim, apt-daily, e2scrub_all, sysstat, and so on). Nothing dumps Postgres and nothing snapshots Garage. The commands in Back up Postgres and Back up the Garage bucket are verified to work, but nothing runs them on a schedule, and nothing copies the result off the host. A backup that lives on the machine it backs up is not a backup.
  4. Fix the Data Protection key ring, per the defect above. Do this before storing real provider keys, because it is a configuration change now and a migration later.
  5. Build a rollback mechanism. This is the largest structural gap. See below.

The rollback gap​

This deployment has no rollback. There is no .deploy.env, no .deploy.prev, no systemd unit and no image pin. Every deploy is a git pull plus a from-source rebuild, and the previous images are overwritten in place.

By contrast the repository does ship a rollback, in infra/compose/, and it is unused here:

infra/compose/ (designed)This deployment (live)
ImagesPrebuilt GHCR, pinned to IMAGE_TAG, a commit SHA, cosign-signedBuilt from source on the host, tagged :latest / :local
Build on the VMImpossible: build: !reset null removes the build contextThe only way images exist
Previous release.deploy.prev holds the prior SHAOverwritten; gone
Rollbackhost.sh deploy a known-good SHANone: rebuild from whatever the branch now says
Supervisionreelbolt.service systemd unitDocker unless-stopped restart policy only
Env validationhost.sh check-env gates every deployNone: a typo surfaces at runtime
Schema safetyMigrations only move forward, so a rollback restores code not schema; "write them expand/contract"Same constraint, no process around it

The practical consequence at 3am: if a deploy is bad, there is nothing to go back to. You can git checkout an older commit and rebuild, but on the media plane you must first solve the missing deploy key, and on both planes the rebuild is the slow, failure-prone part. Database migrations compound it: the Inference API and the engine migrate on startup and only move forward, so rolling code back does not roll the schema back.

The fix is to adopt infra/compose/ and host.sh deploy <sha> rather than to invent something new. It was written for exactly this topology. Until then, treat every deploy as irreversible and take a Postgres dump before up -d --build.