Two-box production deployment
Operational reference for the live two-box ReelBolt deployment: what runs on each machine, why the split is where it is, how to bring it up, how to operate it at 3am, and what is still broken.
Every value here was read off the running boxes rather than from a plan. Where the deployment and an older description disagree, the box wins, and the disagreement is called out inline.
This is not the deployment path the repository ships. infra/compose/ (docker-compose.cloud.yml,
host.sh, env.*.example) is the designed cloud path: SHA-pinned cosign-signed GHCR images, a
reelbolt.service systemd unit, and a .deploy.env / .deploy.prev rollback. This deployment uses
none of that. It builds every image from source on the host and has no rollback mechanism at all. See
What production would still require.
Verified 2026-10-08 against commit b1900986 on both hosts.
The two boxes at a glance
| Control plane | Media plane | |
|---|---|---|
| Hostname | vm-reelbolt-control-prod-weu | vm-reelbolt-media-prod-weu |
| Public IP | 2.29.46.22 | 65.109.128.5 |
| Private IP (Hetzner) | 10.0.0.2 | 10.0.0.3 |
| SSH | ssh -i ~/.ssh/reelbolt_deploy [email protected] | ssh -i ~/.ssh/reelbolt_deploy [email protected] |
| Architecture | x86_64 (amd64) | aarch64 (arm64) |
| CPU | 4 cores | 8 cores |
| RAM | 7751 MiB (7.6 GiB) | 15578 MiB (15.2 GiB) |
| Disk | 75 G, 11 G used (14%) | 150 G, 8.6 G used (6%) |
| OS / kernel | Ubuntu 24.04.1 LTS, kernel 6.8.0-52-generic | Ubuntu 24.04.1 LTS, kernel 6.8.0-52-generic |
| Docker / Compose | 27.5.1 / v2.32.4 | 27.5.1 / v2.32.4 |
| Code | /opt/reelbolt, branch saas-launch | /opt/reelbolt, branch saas-launch |
| Running containers | 11 | 3 |
| Local images | 12 (all amd64) | 5 (all arm64) |
.env | 127 keys, mode 600 root:root | 56 keys, mode 600 root:root |
Both hosts are checked out at b1900986. They are two commits behind the local saas-launch branch
(396bee85): the boxes have not been updated since the legal-placeholder commit.
The two machines talk only over the Hetzner private network (10.0.0.0/16). The media box has no inbound public port except sshd on 22.
Why the split is where it is
The split is not a load-balancing decision. It is a blast-radius decision, and three constraints force it.
Data services and the front door stay on x86
Postgres, Garage (S3), RabbitMQ and Qdrant hold every user's data. The render stack runs untrusted, model-authored code: an agent writes Remotion and ffmpeg workloads that execute in a container. Keeping those two things on different machines means a sandbox escape lands on a host that holds no database.
The front door (nginx) lives with the data because that is where the TLS terminator, the ACME client and the single public 80/443 already are, and because every backend it proxies to other than the engine is on the same host.
The render stack runs on the ARM box because it builds there natively
The media box is aarch64. Every image it runs is built on that box from source; nothing is pulled. This is not a preference, it is the only way those images exist at all. See the architecture constraint.
Only the sandbox-executor holds the Docker socket
sandbox-executor mounts /var/run/docker.sock because it starts the sandbox containers that run
untrusted render workloads. That mount is host root. It is the single most dangerous mount in the
deployment, and it exists on exactly one container on one host, a host that holds no database.
Verified: on the control plane no container mounts the Docker socket at all (the autoheal service,
which would, is pinned behind a never-activated profile). On the media plane exactly one does:
MOUNTS-DOCKER-SOCK: /reelbolt-sandbox-executor-1
/var/run/docker.sock -> /var/run/docker.sock
/var/lib/reelbolt/sandboxes -> /var/lib/reelbolt/sandboxes
What runs where
Control plane, 10.0.0.2 (11 services)
cd /opt/reelbolt
docker compose -f docker-compose.yml -f docker-compose.control.yml up -d
| Service | Image | Arch | Published on |
|---|---|---|---|
postgres | postgres:16-alpine | amd64 | 127.0.0.1:5432, 10.0.0.2:5432 |
garage | dxflrs/garage:v2.4.1 | amd64 | 127.0.0.1:9000, 10.0.0.2:9000 |
rabbitmq | rabbitmq:3-management-alpine | amd64 | 10.0.0.2:5672, mgmt 127.0.0.1:15672 |
qdrant | qdrant/qdrant:latest | amd64 | 10.0.0.2:6333, 10.0.0.2:6334 |
inference | reelbolt-inference:latest (built) | amd64 | 10.0.0.2:8080 |
go-api | reelbolt-go-api:latest (built) | amd64 | none (in-network) |
web | reelbolt-web:latest (built) | amd64 | none (in-network) |
site | reelbolt-site:latest (built) | amd64 | none (in-network) |
docs | reelbolt-docs:latest (built) | amd64 | none (in-network) |
mcp | reelbolt-mcp:latest (built) | amd64 | none (in-network) |
nginx | reelbolt-nginx:latest (built) | amd64 | 0.0.0.0:80, 0.0.0.0:443 |
Also present: alpine:3 (amd64), pulled by the build stages.
Note the container naming. postgres sets an explicit container_name, so it is reelbolt-postgres;
every other service is Compose-generated as reelbolt-<service>-1. Use
docker compose ... exec <service> rather than hardcoding container names.
Nothing is published on the public IP except nginx. The data services bind CONTROL_PRIVATE_IP
(10.0.0.2) and loopback. The override uses Compose's :? form, so an unset value fails loudly instead
of silently binding 0.0.0.0.
Media plane, 10.0.0.3 (4 defined, 3 running)
cd /opt/reelbolt
docker compose -f docker-compose.yml -f docker-compose.media.yml up -d
| Service | Image | Arch | Status |
|---|---|---|---|
workflow-engine | reelbolt-workflow-engine:latest (built) | arm64 | Up, healthy. Publishes 10.0.0.3:8080 |
sandbox-executor | reelbolt-sandbox-executor:local (built) | arm64 | Up, healthy. Mounts the Docker socket |
runner-gateway | reelbolt-runner-gateway:local (built) | arm64 | Up, healthy. In-network 8090 |
sandbox-runtime | reelbolt-sandbox-runtime:local (built) | arm64 | Exited (0), see below |
SERVICE STATUS
runner-gateway Up 8 minutes (healthy)
sandbox-executor Up 9 minutes (healthy)
sandbox-runtime Exited (0) 4 minutes ago
workflow-engine Up 8 minutes (healthy)
sandbox-runtime is not a failure. It is a deliberate build-and-exit service: it builds the
image that untrusted render code runs inside, then stops. Its own compose comment says so:
"Build-and-exit: produces the image untrusted, model-authored render code runs in." An Exited (0) for
this service is the healthy state, and docker compose ps omitting it from the running list is expected.
Do not "fix" it.
Disagreement with the brief. The brief describes four running services on the media plane. Four are defined; only three run.
sandbox-runtimeis a one-shot build stage.
The image it produces is what the executor launches:
SANDBOX_IMAGE=reelbolt-sandbox-runtime:local
The profiling mechanism that keeps the hosts honest
Neither host edits docker-compose.yml. Both add a pure override layer, and each pins the other
plane's services behind a never-activated profile:
- Control pins
workflow-engine,sandbox-runtime,sandbox-executorandrunner-gatewaybehind profilemedia-plane. - Media pins all eleven control services behind profile
media-disabled. - Both pin
whisper,embeddingsandautohealoff. Those have no profile in the base file, so without pinning them they would start.
The media file moves its control-plane services behind a profile, quoting its own comment, "rather than
deleted, so a bare docker compose up on this host can never start a second, split-brain
Postgres/RabbitMQ/Garage or a second API." That is the guardrail: always pass both -f flags.
Bring-up from bare metal
The control plane has a one-shot script; the media plane does not.
1. Both hosts: prerequisites
- Ubuntu 24.04, Docker 27.5.1, Compose v2.32.4.
- The deployment key
~/.ssh/reelbolt_deployauthorized in/root/.ssh/authorized_keys. - The Hetzner private network up, so 10.0.0.2 and 10.0.0.3 are mutually reachable.
/opt/reelbolton branchsaas-launch.
2. Control plane
cd /opt/reelbolt
# .env must exist first: mode 600, see the .env contract below
./bringup.sh
bringup.sh (15 lines, idempotent, re-runnable) is:
COMPOSE="docker compose -f docker-compose.yml -f docker-compose.control.yml"
$COMPOSE build
$COMPOSE up -d
$COMPOSE ps
It is untracked in git: it exists only on the box. Build the control plane first. The Inference API owns the schema and migrates on startup, and the engine on the other host migrates after it.
3. Media plane
There is no bringup.sh here. Run the two commands by hand:
cd /opt/reelbolt
docker compose -f docker-compose.yml -f docker-compose.media.yml build
docker compose -f docker-compose.yml -f docker-compose.media.yml up -d
Build here, before up. The build is what produces the arm64 images; there is nothing to pull.
4. The media box has no GitHub credential, so git pull fails there
Verified on the box:
cd /opt/reelbolt && git fetch --dry-run
# fatal: could not read Username for 'https://github.com': No such device or address
The remote is HTTPS (https://github.com/vecchiotom/ReelBolt.git), there is no credential helper
configured, and /root/.ssh contains only authorized_keys and known_hosts, with no deploy key.
The code reached that box out of band.
Workaround (what was actually done): transfer the tree from the control plane. From the control box:
cd /opt/reelbolt && git bundle create /tmp/reelbolt.bundle saas-launch
Then on the media box:
cd /opt/reelbolt
git fetch /tmp/reelbolt.bundle saas-launch:refs/remotes/bundle/saas-launch
git merge refs/remotes/bundle/saas-launch
The real fix is a read-only deploy key. A GitHub deploy key scoped to this repository, installed in
the media box's /root/.ssh/, plus an SSH remote ([email protected]:vecchiotom/ReelBolt.git). See
What still needs the owner. Until then the media box cannot be deployed
to independently, and any update must be pushed from the control plane.
Disagreement with the brief. The brief says both boxes run code at
/opt/reelbolton branchsaas-launchwithout qualification. True, but the media checkout is not deployable on its own, and neither checkout is clean:bringup.shanddocker-compose.control.ymlare untracked on the control box, anddocker-compose.media.ymlis untracked on the media box. Those override files exist nowhere in git. Back them up separately:git clean -fdxwould destroy the deployment.
5. What starts the stack after a reboot
There is no systemd unit. Verified on both hosts: no reelbolt.service. What brings the
stack back after a reboot is Docker's own restart policy — every container is unless-stopped, and
docker.service is enabled, so the daemon restarts them all by itself.
The consequence for the 3am reader: containers come back by themselves; new code does not. A reboot restores the running stack, not a deployed one.
Since 2026-10-09 there is a release mechanism and a rollback target: infra/scripts/release.sh
records the live commit in .deploy.sha and the outgoing one in .deploy.prev. See "Releases and
rollback" below. (This replaces bringup.sh, which ran compose build && compose up -d and left no
record of what had been running, so a bad release could not be undone. bringup.sh is no longer
tracked and should not be used.)
Releases and rollback
A release is a git commit. infra/scripts/release.sh runs on the box and is the only thing that
should ever start this stack by hand.
# On either box, from /opt/reelbolt:
REELBOLT_ROLE=control bash infra/scripts/release.sh deploy # release the current checkout
REELBOLT_ROLE=control bash infra/scripts/release.sh deploy <sha> # fetch and release a specific commit
REELBOLT_ROLE=control bash infra/scripts/release.sh rollback # back to .deploy.prev
REELBOLT_ROLE=control bash infra/scripts/release.sh status
REELBOLT_ROLE=control bash infra/scripts/release.sh prune 5 # keep the 5 newest releases
REELBOLT_ROLE is control or media and selects both the override layer
(docker-compose.control.yml / docker-compose.media.yml) and the health checks: control probes
https://127.0.0.1/{health,api/v1/health,api/v1/workflow-engine/health} through nginx, so a green
release means the front door and everything behind it answer; media waits on the engine's own
container healthcheck.
How a release is recorded
.deploy.sha the release that is live now
.deploy.prev the release that was live before it
Deploying tags every locally-built image as <repository>:<sha>, so the previous release's images
survive the next one. rollback re-points the reference compose starts each container from at the
previous tag and restarts — no rebuild, so it takes as long as a restart. A failed health check
performs exactly that rollback automatically and then exits non-zero.
Two things about image names, because both have already caused a failed release:
- The reference is not always
<project>-<service>. The base compose pinsreelbolt-sandbox-runtime:local,reelbolt-sandbox-executor:localandreelbolt-runner-gateway:local; the rest are left for compose to name, givingreelbolt-<service>:latest.release.shreads the composed config rather than assuming, and keeps the repository and the tag separate — using the tagged reference as the base for the release tag producesreelbolt-workflow-engine:latest:<sha>, which Docker rejects outright. - Schema is not rolled back. Migrations only move forward, so a rollback restores code, not the database. Write migrations expand/contract.
Both boxes hold a read-only GitHub deploy key at /root/.ssh/reelbolt_github (registered on the
repo as "two-box deploy (read-only)"), with GitHub's host key pinned in /root/.ssh/known_hosts, so
deploy <sha> can fetch a commit that is not on the box yet.
Retiring Garage
Object storage moved to Cloudflare R2 (2026-10-09) and Garage is pinned to the cloud-disabled
profile so compose up -d cannot resurrect it. The reelbolt_garagedata/reelbolt_garagemeta
volumes are kept deliberately, so the migration stays reversible and the old objects remain readable.
go-api's depends_on had to be overridden to drop Garage: a depends_on naming a service that no
active profile defines invalidates the entire compose project, with the unhelpful message
service "go-api" depends on undefined service "garage".
The config keys are still spelled MinIO__* / MINIO_*; that is historical. What matters for R2 is
MINIO_REGION=auto (Garage enforced garage) and path-style addressing.
Certificate renewal
Caddy obtains and renews the certificate; nginx only reads it. The one non-obvious requirement — the
Host header on the ACME challenge location, without which renewal fails silently for ~60 days — is
in tls.md under "Renewals". Read that before touching nginx/http-server-redirect.conf.
The .env contract
Both hosts read /opt/reelbolt/.env, mode 600 root:root. These files hold every credential the
deployment has. Never print their contents, never paste them into a ticket, never commit them. Every
value in this section is referenced by variable name only.
- Control: 127 keys, 35
# PASTE:markers. - Media: 56 keys, 11
# PASTE:markers.
Variables that must be byte-identical across the two hosts
This is the part that bites. The two services share a Postgres database, a RabbitMQ broker and an object store, and the Go API issues JWTs that the engine validates. Each value below was compared across the two boxes by hashing it and comparing digests; every one is confirmed identical. Changing any of them on one host only breaks the deployment.
| Variable | Why it must match |
|---|---|
JWT_SIGNING_KEY | The Go API signs; the Inference API and the engine verify. A mismatch invalidates every session token and every engine-to-API call. |
POSTGRES_USER, POSTGRES_PASSWORD, POSTGRES_DB | One database, two hosts. |
RABBITMQ_USER, RABBITMQ_PASSWORD | The engine consumes from the control plane's broker. |
MINIO_ACCESS_KEY, MINIO_SECRET_KEY, MINIO_BUCKET | One Garage bucket, written by both hosts. |
SANDBOX_API_TOKEN | The engine authenticates to the sandbox executor with it. |
RUNNER_INTERNAL_TOKEN | Engine to runner-gateway internal auth. |
RUNNER_JOB_TOKEN_KEY | Signs runner job tokens; minted on one side, verified on the other. |
CONTROL_PRIVATE_IP (10.0.0.2) | Each host needs the other's address; both hold both. |
ENGINE_PRIVATE_IP (10.0.0.3) | As above. |
TENANCY_MODE | Currently Cloud on both. A split here makes the two services disagree about tenancy. |
Present on the control plane but not on the media plane
INTERNAL_API_TOKEN: set on control, absent on media. The media.envdoes not define it at all. Nothing on the media plane currently reads it, so this is not an active fault, but it is an asymmetry to know about before wiring anything new that expects it.- The entire front-door and third-party block:
DOMAIN,PUBLIC_BASE_URL,SIGNUP_ENABLED,OAUTH_*,SMTP_*,PADDLE_*,TURNSTILE_*,PLATFORM_ASR_*,PLATFORM_EMBEDDING_*and the rest.
Present on the media plane but not on the control plane
DATABASE_CONNECTION_STRING: the .NET connection string, pointing atHost=10.0.0.2;Port=5432.DATABASE_URL: points at 10.0.0.2.MINIO_ENDPOINT,MINIO_REGION,MINIO_FORCE_PATH_STYLE: the control plane omits these and relies on the base compose defaults; the media override requires them.DATA_PROTECTION_STORE,AWS_REGION: see the defect below. Both are currently inert.- The media-only render tuning keys:
VIDEO_FFMPEG_THREADS,OTEL_*.
Both hosts: DATABASE_URL and DATABASE_CONNECTION_STRING embed the Postgres password in plaintext
This is a real finding, stated without the value. Every .env also carries POSTGRES_PASSWORD
separately, so the same credential appears twice in the control file and three times in the media
file.
Rotating the Postgres password therefore means editing five places across two hosts, and missing one
produces a failure that looks like a network problem. Change POSTGRES_PASSWORD, DATABASE_URL and
DATABASE_CONNECTION_STRING together on the media box, and POSTGRES_PASSWORD and DATABASE_URL on the
control box. For reference, the connection strings are shaped like:
DATABASE_URL=postgres://<user>:<REDACTED>@10.0.0.2:5432/reelforge?sslmode=disable
DATABASE_CONNECTION_STRING=Host=10.0.0.2;Port=5432;Database=reelforge;Username=<user>;Password=<REDACTED>
The PASTE markers
Both files ship from .env.example with every third-party credential blank and marked # PASTE:.
- Control: 35 markers. Azure OpenAI (endpoint and key), Anthropic (Console key and subscription token), Google AI Studio and Google Cloud keys, DeepSeek key, the hosted ASR provider block (name, endpoint, model, key), the hosted embedding block (name, endpoint, model, key), SMTP (host, username, password, from), Turnstile (secret and site key), Google and GitHub OAuth (id and secret pairs), hosted Qdrant URL and key, ACME e-mail, GA4 measurement id, contact webhook URL, and the five Paddle values (API key, webhook secret, three top-up price ids).
- Media: 11 markers. Five are copy-from-control markers (
JWT_SIGNING_KEY, RabbitMQ,DATABASE_CONNECTION_STRING,DATABASE_URL, MinIO) and those are already filled. The remaining six are provider credentials (Anthropic, Google, DeepSeek, Azure OpenAI) plus an OTEL authorization header, and are still blank.
A blank here is not always a fault. SIGNUP_ENABLED=false and BILLING_ENABLED=false mean the Paddle and
Turnstile blocks are legitimately empty today. But PLATFORM_ASR_NAME and PLATFORM_EMBEDDING_NAME are
both blank on the control plane, which means no transcription and no embedding provider is configured
at all: the local whisper and embeddings containers are pinned off, so both capabilities currently
have no source. Any workflow step needing ASR or semantic file search fails until an owner pastes one.
The architecture constraint: aarch64 vs x86_64
The control plane is x86_64 (amd64). The media plane is aarch64 (arm64). They cannot share images. This is the constraint that shapes every deploy decision on the media box.
Verified: every image on the media box is arm64, and every image on the control plane is amd64.
for i in $(docker images --format '{{.Repository}}:{{.Tag}}' | sort -u); do
echo "$i -> $(docker image inspect --format '{{.Architecture}}/{{.Os}}' "$i")"
done
# media plane (aarch64 host)
alpine:3 arm64/linux
reelbolt-runner-gateway:local arm64/linux
reelbolt-sandbox-executor:local arm64/linux
reelbolt-sandbox-runtime:local arm64/linux
reelbolt-workflow-engine:latest arm64/linux
# control plane (x86_64 host)
alpine:3 amd64/linux
dxflrs/garage:v2.4.1 amd64/linux
postgres:16-alpine amd64/linux
qdrant/qdrant:latest amd64/linux
rabbitmq:3-management-alpine amd64/linux
reelbolt-docs:latest amd64/linux
reelbolt-go-api:latest amd64/linux
reelbolt-inference:latest amd64/linux
reelbolt-mcp:latest amd64/linux
reelbolt-nginx:latest amd64/linux
reelbolt-site:latest amd64/linux
reelbolt-web:latest amd64/linux
Every image the media box runs is built on the media box
The media override states it directly: "The images are built HERE, on this ARM64 box, instead of pulled from GHCR. A plain docker build on aarch64 produces an arm64 image, which is what the control plane (x86_64) must never receive. Nothing amd64 is pulled."
So on the media plane the sequence is always build then up -d. There is no registry to fall back on
and no IMAGE_TAG pin.
Verify the architecture of anything you rebuild
docker image inspect --format '{{.Architecture}}' reelbolt-workflow-engine:latest
# must print: arm64
Check all four after any rebuild:
for i in reelbolt-workflow-engine:latest reelbolt-sandbox-executor:local reelbolt-sandbox-runtime:local reelbolt-runner-gateway:local; do
printf '%-42s %s\n' "$i" "$(docker image inspect --format '{{.Architecture}}' "$i")"
done
Why pulling an amd64 image there would fail, or silently emulate
- Fail: the kernel cannot execute amd64 ELF binaries. You get
exec format errorat container start, often deep inside a healthcheck loop, with the container restarting repeatedly. - Silently emulate: if binfmt_misc or QEMU handlers are registered on the host, Docker will happily run
the amd64 image under emulation. Nothing errors. It is simply wrong: an arm64 host emulating amd64 runs
an ffmpeg or Remotion render several times slower, at which point
VIDEO_MAX_CONCURRENT_JOBS=2oversubscribes 8 real cores and renders start timing out. This is the dangerous case, because the symptom is "the box got slow", not "the box is broken".
The sharp edge: SANDBOX_IMAGE=reelbolt-sandbox-runtime:local is what sandbox-executor launches for
every untrusted render. If a rebuild ever produces an amd64 reelbolt-sandbox-runtime:local, every sandbox
container is wrong-arch. Sandbox containers are created on demand, so this never shows up in
docker compose ps. It shows up as failed renders.
Day-2 operations
All commands run from /opt/reelbolt with the correct pair of -f flags. Define them once:
# control plane
alias rbc='docker compose -f docker-compose.yml -f docker-compose.control.yml'
# media plane
alias rbm='docker compose -f docker-compose.yml -f docker-compose.media.yml'
Restart one service
# control plane
cd /opt/reelbolt
docker compose -f docker-compose.yml -f docker-compose.control.yml restart go-api
# media plane
cd /opt/reelbolt
docker compose -f docker-compose.yml -f docker-compose.media.yml restart workflow-engine
On the control plane a restart can cascade, and this is expected. nginx declares depends_on with
restart: true for site, web, go-api and inference. Restarting any of those also restarts
nginx, a brief blip on the public front door. Verified:
$ docker compose -f docker-compose.yml -f docker-compose.control.yml restart go-api
Container reelbolt-go-api-1 Restarting
Container reelbolt-go-api-1 Started
Container reelbolt-nginx-1 Restarting <-- cascade
Container reelbolt-nginx-1 Started
The media plane does not cascade: restarting workflow-engine leaves runner-gateway and
sandbox-executor untouched.
restart does not rebuild, which is the common case. To pick up new code, use up -d --build <service>.
Read logs
# follow one service on the control plane
docker compose -f docker-compose.yml -f docker-compose.control.yml logs -f --tail=200 inference
# media engine
docker compose -f docker-compose.yml -f docker-compose.media.yml logs -f --tail=200 workflow-engine
# since a timestamp
docker compose -f docker-compose.yml -f docker-compose.control.yml logs --since 30m go-api
By container name when you already know it. Remember reelbolt-postgres is the one exception to the
-1 suffix rule:
docker logs --tail=200 reelbolt-postgres
docker logs --tail=200 reelbolt-workflow-engine-1
The engine is the busiest log in the system. Start there when a workflow misbehaves.
Check health
The authoritative source is the container healthchecks themselves. Read what the box actually runs rather than guessing a path:
for c in $(docker ps -q); do
n=$(docker inspect -f '{{.Name}}' $c | sed 's|^/||')
echo "$n :: $(docker inspect -f '{{.State.Health.Status}}' $c 2>/dev/null || echo no-healthcheck)"
done
Real healthcheck endpoints on this deployment:
| Service | Healthcheck |
|---|---|
nginx | https://127.0.0.1/health and https://127.0.0.1/api/v1/health |
go-api | http://localhost:8080/health |
inference | http://localhost:8080/api/v1/health |
web | http://127.0.0.1:3000/app/login |
site | http://127.0.0.1:3000/ |
docs | http://127.0.0.1/healthz |
mcp | http://localhost:3002/health |
postgres | pg_isready -U postgres -d reelforge |
garage | /garage status |
rabbitmq | rabbitmq-diagnostics -q check_running |
workflow-engine (media) | http://localhost:8080/api/v1/workflow-engine/health |
sandbox-executor (media) | http://localhost:8080/health |
runner-gateway (media) | http://localhost:8090/healthz |
qdrant has no healthcheck. It shows as Up with no health state, which is normal.
The engine's health endpoint is at /api/v1/workflow-engine/health, not /health. The bare paths
404. Verified through the private network from the control plane:
curl -s http://10.0.0.3:8080/api/v1/workflow-engine/health
# {"status":"healthy","service":"workflow-engine","timestamp":"..."}
Front-door check. Note -k: the certificate is self-signed.
curl -sk -o /dev/null -w '%{http_code}' https://2.29.46.22/ ; echo
curl -sk -o /dev/null -w '%{http_code}' https://2.29.46.22/api/v1/health ; echo
curl -sk -o /dev/null -w '%{http_code}' https://2.29.46.22/api/v1/workflow-engine/health ; echo
Back up Postgres
Postgres holds all application state: users, organizations, projects, workflows, executions, metering, billing and the schema itself. This is the only backup that really matters.
# on the control plane
docker exec reelbolt-postgres pg_dump -U postgres -d reelforge -Fc > /root/reelforge-$(date -u +%Y%m%dT%H%M%SZ).dump
Verified working: pg_dump (PostgreSQL) 16.15, database size 10207 kB, a plain-SQL dump of about
302 KB. Small, but the billing and metering ledgers cannot be reconstructed from anywhere else.
Restore:
docker exec -i reelbolt-postgres pg_restore -U postgres -d reelforge --clean --if-exists < /root/reelforge-YYYYMMDDTHHMMSSZ.dump
pg_dump alone is not sufficient here. See the Data Protection defect below: restoring the database
without the matching key ring leaves every encrypted provider key permanently unreadable.
Back up the Garage bucket
Garage is the S3-compatible object store. The bucket is reelforge.
docker exec reelbolt-garage-1 /garage bucket list
docker exec reelbolt-garage-1 /garage key list
docker exec reelbolt-garage-1 /garage status
Verified: one node 76537915d73d7d9f, capacity 74.8 GiB with 61.7 GiB available, bucket reelforge, one
key. Object data is currently tiny, about 8 KB of volume data, because the stack holds no production media
yet.
The Garage container has no shell. sh is not in the image, so docker exec ... sh -c fails with
exec: "sh": executable file not found. Use the /garage binary directly, as above.
Garage provides no backup subcommand. Back it up at the filesystem level, with the container stopped so nothing is mid-write:
# on the control plane
docker compose -f docker-compose.yml -f docker-compose.control.yml stop garage
tar czf /root/garage-$(date -u +%Y%m%dT%H%M%SZ).tar.gz -C /var/lib/docker/volumes reelbolt_garagedata reelbolt_garagemeta
docker compose -f docker-compose.yml -f docker-compose.control.yml start garage
The two volumes are confirmed as:
/var/lib/docker/volumes/reelbolt_garagedata/_datamounted at container/data(objects)/var/lib/docker/volumes/reelbolt_garagemeta/_datamounted at container/meta(metadata)
You need both. Restoring /data without /meta gives you blocks with no index, and
garage repair blocks is an involved recovery, not a quick fix. Garage also stores the cluster
rpc_secret in /opt/reelbolt/garage/garage.toml. Back that up too, or the restored node has a different
identity.
Verify the cross-box link
This is the single most important check when something "cannot reach the database". Run all of it from the control plane.
1. RabbitMQ: which hosts are connected to the broker. The media engine's connections appear with
peer_host 10.0.0.3:
docker exec reelbolt-rabbitmq-1 rabbitmqctl list_connections name peer_host
Observed:
name peer_host
172.18.0.8:44112 -> 172.18.0.5:5672 172.18.0.8 <- in-network (control)
172.18.0.9:59470 -> 172.18.0.5:5672 172.18.0.9 <- in-network (control)
10.0.0.3:59256 -> 172.18.0.5:5672 10.0.0.3 <- MEDIA ENGINE
10.0.0.3:42154 -> 172.18.0.5:5672 10.0.0.3 <- MEDIA ENGINE
Two connections from 10.0.0.3 is the proof the cross-box link is up. If they are missing, the engine is down, the private network is broken, or the credentials diverge.
Queue state. The engine is the consumer:
docker exec reelbolt-rabbitmq-1 rabbitmqctl list_queues name messages consumers
Observed: ProjectFileIndexing, workflow-execution, workflow-stop-requests and
go-api-workflow-events, each with a consumer and 0 messages backed up.
2. Postgres: which hosts hold connections. The media engine's pool appears as 10.0.0.3:
docker exec reelbolt-postgres psql -U postgres -d reelforge -c "select client_addr, usename, datname, state, count(*) from pg_stat_activity where backend_type = 'client backend' group by 1,2,3,4 order by 1;"
Observed:
client_addr | usename | datname | state | count
-------------+----------+-----------+--------+-------
10.0.0.3 | postgres | reelforge | idle | 5 <- MEDIA ENGINE POOL
172.18.0.8 | postgres | reelforge | idle | 1 <- inference (in-network)
172.18.0.9 | postgres | reelforge | idle | 1 <- go-api (in-network)
| postgres | reelforge | active | 1 <- this query
Five idle connections from 10.0.0.3 is the proof the engine reaches the database. The filter must be
backend_type = 'client backend'; with an underscore the query errors out.
3. Raw TCP reachability. If the two checks above fail, this says whether it is the network or the application:
# from the MEDIA box, toward the control plane
for hp in 10.0.0.2:5432 10.0.0.2:5672 10.0.0.2:9000 10.0.0.2:6333 10.0.0.2:8080; do
h=${hp%:*}; port=${hp#*:}
if timeout 5 bash -c "cat < /dev/null > /dev/tcp/$h/$port" 2>/dev/null; then echo "OPEN $hp"; else echo "CLOSED $hp"; fi
done
All five are currently OPEN: Postgres, RabbitMQ, Garage, Qdrant REST and the Inference API.
4. Confirm the media box still has no inbound public port:
ss -tln | grep -v '127.0.0.1' | grep -v '10.0.0.3'
Expected: only sshd.
LISTEN 0 4096 *:22 *:*
Known defects and their current state
Re-check these before acting on them. At the time of writing two of the three were still live, and another agent may have fixed one since. Every claim below is what was observed on the boxes on 2026-10-08, not what was planned.
1. The Data Protection key ring is per-host: STILL BROKEN
Status: open. Currently latent, because no encrypted keys are stored yet, but guaranteed to bite the first time an admin types an API key into the UI.
The Inference API encrypts provider API keys at rest with ASP.NET Core Data Protection (ISecretProtector).
The engine decrypts what the API wrote, so both services must share one key ring. They do not.
Each host has its own dpkeys volume, and the two volumes contain different keys:
# control plane
reelbolt_dpkeys/key-7aa7e137-1680-4bea-b7c0-f35fc4270a76.xml
# media plane
reelbolt_dpkeys/key-f9792ecb-2b94-482e-b7fe-823c58e5e5bf.xml
Two different key-ring files on two hosts. Neither can read the other's ciphertext.
Why it is latent right now: the inference_providers table is empty, 0 rows on the deployment.
Nothing has been encrypted yet. Keys whose value is the literal env: sentinel are unaffected, because
they are never persisted.
The trap: the failure mode is not a startup error. It appears at first use: a provider row saved on
the control plane, then a workflow step on the media engine that tries to decrypt it and fails. Nothing in
docker compose ps will look wrong.
The evidence that the code is wired correctly but the configuration is not: the base
docker-compose.yml sets only DataProtection__KeysPath: /keys (inference at line 331, engine at line
443). It sets no DataProtection__Store. The repository's own cloud override
infra/compose/docker-compose.cloud.yml does set DataProtection__Store: ${DATA_PROTECTION_STORE:-Database},
with the comment: "a per-host /keys directory would leave the other VM unable to read keys this one
wrote." The designed cloud path solves this; this deployment does not use it.
A red herring to be aware of: the media .env contains DATA_PROTECTION_STORE=FileSystem. That
variable is inert; nothing reads it. DATA_PROTECTION_STORE appears nowhere in docker-compose.yml,
and the engine container's environment contains only DataProtection__KeysPath=/keys. Editing that line
changes nothing. AWS_REGION in the media .env is likewise a leftover.
What the fix requires: set DataProtection__Store=Database on both the Inference API and the
engine, plus a shared wrapper key: either DATA_PROTECTION_KMS_LOCAL_KEY (a base64 32-byte value,
identical on both hosts) or AWS KMS via DATA_PROTECTION_KMS_KEY_ID plus credentials. Then restart both
services. See docs/secrets.md and infra/compose/docker-compose.cloud.yml for the exact keys. Do this
before anyone stores a real provider key, because migrating an existing per-host ring is a manual
re-wrap, not an automatic one.
2. Presigned URLs point at a private address: STILL SET, currently harmless
Status: the misconfiguration is present; its impact is latent because the compute fabric is off.
MINIO_PUBLIC_ENDPOINT=http://10.0.0.2:9000 on both hosts: the private Hetzner address.
MinIO__PublicEndpoint is the host presigned URLs are signed against. SigV4 signs the Host header,
so this value must be the address the consumer of the URL will actually connect to.
docs/cloud-storage.md states the rule: "it must be the host a runner or browser really connects to."
10.0.0.2 is unroutable from anywhere outside the private network.
Why it is harmless today:
COMPUTE_MODE=Localon both hosts. The only consumers of engine-minted presigned URLs are the runner-fabric paths (SandboxWorkspaceJournal,RemoteSandboxSession,OutputVerifier). With the fabric off they are never exercised.- The browser never receives a presigned URL. Downloads are streamed through the Inference API
(
RangedObjectStreamer.StreamAsyncover the API's own in-network S3 client), not redirected to the object store. So this defect cannot break a dashboard download today.
When it will bite: the moment COMPUTE_MODE=Fabric is enabled, or a desktop runner
(docs/desktop-runner.md), a machine outside the private network, receives a manifest. Every presigned
GET and PUT it is handed will name an address it cannot reach. The failure is a timeout or a connection
refused inside the sandbox, not an auth error, which makes it look like a network problem.
A correctness note on the current value: because the runner and the sandbox containers all live on the media box, they can reach 10.0.0.2:9000, so a Fabric rollout would appear to work and then fail for exactly the first off-network consumer. If a real domain is provisioned, this value must change to the public HTTPS endpoint of the object store at the same time.
3. Tenancy mode: FIXED, verified Cloud on both hosts
Status: correct. This is what the box shows, not what was planned.
TENANCY_MODE=Cloud is set in the .env of both hosts, is passed through to go-api by the control
override, and is live in the running container:
docker exec reelbolt-go-api-1 printenv TENANCY_MODE
# Cloud
Corroborated by the data. The tenancy shape is the Cloud one, with no shared Default organization:
docker exec reelbolt-postgres psql -U postgres -d reelforge -c "select kind, count(*) from organizations group by kind;"
kind | count
----------+-------
Personal | 1
Team | 1
Platform | 1
A SelfHosted deployment would put users in a shared Default organization. There is none.
Why this variable must be passed explicitly. From the control override: the base compose file never
sets TENANCY_MODE, so "a cloud deployment would silently run SelfHosted (users join a shared Default
org)." If you ever rewrite the override, keep this line. Its absence is silent, and the symptom is a
permissions model that is not the one you think you deployed.
Related and also correct: SIGNUP_ENABLED=false and NEXT_PUBLIC_SIGNUP_ENABLED=false.
What still needs the owner
Four things cannot be done from inside the deployment, because each needs an external credential, a DNS record, or a decision that is not an engineer's to make.
1. A real domain: the stack runs on a bare IP with a self-signed certificate
| Variable | Value |
|---|---|
PUBLIC_BASE_URL | http://2.29.46.22 |
NEXT_PUBLIC_SITE_URL | http://2.29.46.22 |
DOMAIN | empty |
ACME_EMAIL | empty |
FORCE_HTTPS | false |
COOKIE_SECURE | false |
nginx serves a self-signed certificate:
subject=CN = localhost
issuer=CN = localhost
notBefore=Oct 8 22:55:23 2026 GMT
notAfter =Oct 8 22:55:23 2027 GMT
Consequences: browsers show a certificate warning (tls_verify_result=18, self-signed), there is no
trusted TLS, and plain HTTP on port 80 serves the application directly with no redirect to HTTPS.
Disagreement with the brief. The brief says the stack runs on
http://2.29.46.22with a self-signed cert. Both halves are true simultaneously: port 80 and port 443 are both open and both return 200. That is worse than "HTTP only", because credentials can be sent in cleartext to a working endpoint while the HTTPS listener sits there ignored.
Needed: a domain and its A records, DOMAIN and ACME_EMAIL set in .env, and the documented
cut-over in docs/tls.md, which activates the caddy container, already pinned behind the opt-in tls
profile and left untouched for exactly this purpose.
2. Third-party credentials: 35 blank markers on control, 11 on media
Every third-party integration is unconfigured. The full inventory is in the PASTE markers section. The ones that block a launch:
- A transcription (ASR) provider and an embedding provider.
PLATFORM_ASR_NAMEandPLATFORM_EMBEDDING_NAMEare blank, and the localwhisperandembeddingscontainers are pinned off. The deployment currently has no speech-to-text and no semantic search at all. - A chat provider. Every chat key (
ANTHROPIC_API_KEY,GOOGLE_API_KEY,DEEPSEEK_API_KEY,AZURE_OPENAI_API_KEY) is blank, so no agent can run. - SMTP. Password resets and email verification fall back to console logging, so no user can reset a password.
- Paddle (key, webhook secret, three price ids), required before
BILLING_ENABLED=true. - OAuth (Google, GitHub) and Turnstile: optional, currently all disabled.
- ACME e-mail, required for TLS.
3. A read-only deploy key (GitHub)
The media box has no GitHub credential and cannot pull. See the media box has no GitHub credential. Needed: a repository-scoped read-only deploy key, not a personal token and not a write key, installed on the media box, plus an SSH remote. Read-only is the right scope: the media box only ever needs to fetch.
4. TLS
Covered above under the domain. docs/tls.md is the runbook; the caddy container and the nginx
/.well-known/acme-challenge/ location are already wired and waiting.
Security posture and its rationale
Why the media box has no inbound public port
The media box runs untrusted, model-authored code. Its job is to execute code that an AI agent wrote, on behalf of a user, with no human review. Giving that host a public attack surface would be indefensible. It has one listening public port, sshd, and everything else is private:
LISTEN 0 4096 127.0.0.54:53 0.0.0.0:*
LISTEN 0 4096 127.0.0.53%lo:53 0.0.0.0:*
LISTEN 0 4096 10.0.0.3:8080 0.0.0.0:* <- engine, PRIVATE address only
LISTEN 0 4096 *:22 *:* <- the only public inbound
The engine publishes on 10.0.0.3:8080, never 0.0.0.0:8080. The control plane's nginx is the only caller
(UPSTREAM_WORKFLOW_ENGINE=http://10.0.0.3:8080).
Why only the sandbox-executor mounts the Docker socket
sandbox-executor must start containers, so it needs /var/run/docker.sock, which is host root.
Anyone who escapes that container owns the host. The response is not to remove the mount, since the service
cannot work without it, but to bound what the host holds:
- The socket lives on the media box, which holds no database. The database is on the other host.
- The mount is on exactly one container on that host. Verified on both boxes.
- On the control plane nothing mounts the socket at all. The
autohealservice, which would, is pinned behind the never-activatedcloud-disabledprofile, and the override says why: "With it off, NOTHING on this box mounts the Docker socket." - The control plane's
webservice has its 9229 Node inspector publish reset to empty, described in the override as "unauthenticated RCE for anyone who can reach the port."
This is the whole reason the two-box split exists. A sandbox escape on the media box reaches a host with no user data, no database and no object store.
Password authentication is still enabled on both hosts, deliberately
Both hosts currently have:
permitrootlogin yes
pubkeyauthentication yes
passwordauthentication yes
Confirmed in /etc/ssh/sshd_config.d/50-cloud-init.conf on both.
This is a deliberate, temporary state. Key authentication is proven to work: auth.log shows 90
successful publickey logins on the control plane and 61 on the media plane. Password auth was left on as
a fallback and the cut-over was never performed.
The weak point is not the key auth. It is that PermitRootLogin yes combined with password auth means
the internet can attempt to brute-force root, and root's only constraint is PAM. There is no fail2ban
configured on either host.
What production would still require
In priority order:
- Turn off password authentication. Set
PasswordAuthentication noin/etc/ssh/sshd_config.d/50-cloud-init.conf, or betterPermitRootLogin prohibit-password, on both hosts, and validate the key-only path in a second session before closing the first. Key auth is already proven, so this is a five-minute change that removes an unbounded risk. - Terminate real TLS. A domain plus the documented
docs/tls.mdcut-over. Also setCOOKIE_SECURE=trueandFORCE_HTTPS=trueonce HTTPS is the only listener. TodayCOOKIE_SECURE=falsemeans session cookies are sent over plain HTTP. - Configure backups. There are none. Neither host has a root crontab, and every systemd timer on
both hosts is an OS default (
dpkg-db-backup,logrotate,fstrim,apt-daily,e2scrub_all,sysstat, and so on). Nothing dumps Postgres and nothing snapshots Garage. The commands in Back up Postgres and Back up the Garage bucket are verified to work, but nothing runs them on a schedule, and nothing copies the result off the host. A backup that lives on the machine it backs up is not a backup. - Fix the Data Protection key ring, per the defect above. Do this before storing real provider keys, because it is a configuration change now and a migration later.
- Build a rollback mechanism. This is the largest structural gap. See below.
The rollback gap
This deployment has no rollback. There is no .deploy.env, no .deploy.prev, no systemd unit and no
image pin. Every deploy is a git pull plus a from-source rebuild, and the previous images are overwritten
in place.
By contrast the repository does ship a rollback, in infra/compose/, and it is unused here:
infra/compose/ (designed) | This deployment (live) | |
|---|---|---|
| Images | Prebuilt GHCR, pinned to IMAGE_TAG, a commit SHA, cosign-signed | Built from source on the host, tagged :latest / :local |
| Build on the VM | Impossible: build: !reset null removes the build context | The only way images exist |
| Previous release | .deploy.prev holds the prior SHA | Overwritten; gone |
| Rollback | host.sh deploy a known-good SHA | None: rebuild from whatever the branch now says |
| Supervision | reelbolt.service systemd unit | Docker unless-stopped restart policy only |
| Env validation | host.sh check-env gates every deploy | None: a typo surfaces at runtime |
| Schema safety | Migrations only move forward, so a rollback restores code not schema; "write them expand/contract" | Same constraint, no process around it |
The practical consequence at 3am: if a deploy is bad, there is nothing to go back to. You can
git checkout an older commit and rebuild, but on the media plane you must first solve the missing deploy
key, and on both planes the rebuild is the slow, failure-prone part. Database migrations compound it: the
Inference API and the engine migrate on startup and only move forward, so rolling code back does not
roll the schema back.
The fix is to adopt infra/compose/ and host.sh deploy <sha> rather than to invent something new. It was
written for exactly this topology. Until then, treat every deploy as irreversible and take a Postgres
dump before up -d --build.