The CloudPool worker
What the CloudPool worker is
reelbolt-worker is the binary a compute-plane VM runs. It dials the runner
gateway over a WebSocket, authenticates with the pool credential, advertises what
that machine can actually do, and executes the jobs the gateway assigns to it.
It is the first embedder of the runner library in sandbox/pkg/agent: the
protocol, the connection, the lease renewal, the op dedupe cache and the
reconnect backoff all live there, and this binary supplies only the
configuration, the capability probe, the session executors and the process
lifecycle.
It is deliberately not the sandbox-executor. That process is the self-host
HTTP control plane: it listens, the engine calls it, and it holds a bearer token
so nobody else can. A pool worker listens on nothing, is reached by nobody, and
is recognised by the gateway because it dialled out. The two share
sandbox/pkg/container and nothing else.
The spec it implements is the runner protocol.
Enrolling a pool worker (and why the worker cannot enroll itself)
A pool worker with no runner_devices row is refused 4401. The gateway
requires a row with kind='CloudPool' whose name matches the workerName
in the hello, and it refuses rather than inventing a synthetic identity. That is
deliberate: an identity nobody enrolled is an identity nobody can revoke.
The command and the cloud-init fragment exist in this tree — this binary is
D13, with its out-of-band enroll subcommand — but no running deployment
runs them: no cloud provider is configured here, no node boots, and nothing
writes a CloudPool runner_devices row. So today a pool job created through
the internal API is created successfully and then stays Queued, with no
error, however healthy the pool looks. That is a missing data plane (no worker
runs in this tree), not a missing caller, and it is stated plainly rather than
implied. Pairing a desktop runner is a different path and does write a
runner_devices row. What a provider launch (F12) must supply is in "What F12
must supply" below. Enrollment is the command that closes the gap when a
provider exists:
DATABASE_URL=postgres://... REELBOLT_WORKER_NAME=pool-eu-1 reelbolt-worker enroll
enroll is idempotent — the same name always lands on the same row, and a
second boot updates os, arch and runner_version rather than inserting a
rival — and it is safe to run on every boot. infra/worker/cloud-init.yaml
does exactly that, in its own systemd unit before the worker unit.
Three properties are worth knowing before changing it:
- It never clears
revoked_at. A revoked identity stays revoked until an operator passes-reinstate. If enrollment resurrected a burned worker on the next boot, revocation would be a decoration. - It refuses an ambiguous name. The gateway's lookup is
ORDER BY created_at LIMIT 1, so two rows sharing a name mean the worker's identity depends on row age.enrollreports the conflict and writes nothing. - The running worker must not hold the database credential.
reelbolt-worker runrefuses to start withDATABASE_URLset. The enrollment credential and the service credential are different things, and only the first needs write access torunner_devices.
The same refusal is what an operator sees when something goes wrong: a 4401 logs a line naming the worker, the missing row and the exact command that creates it, and exits with code 4 so a restart loop does not paper over it.
A worker name is an identity, and two live processes may not share one. The gateway supersedes by identity (the runner protocol, "the new connection wins and the old one is closed with 4409"), and 4409 is terminal in the agent library, so the displaced process exits and does not fight for the socket. That is the right behaviour and the wrong deployment: give each VM its own name (the hostname default does this) and never run two workers with one enrolled row. It is also why the local smoke kills its own leftovers before it starts: a surviving worker from an interrupted run reconnects to the same port and steals the socket from the run that is starting.
The worker's environment
| Variable | Meaning |
|---|---|
REELBOLT_GATEWAY_URL | wss://runner.example/v1/connect. Required. |
RUNNER_POOL_CREDENTIAL | The pool secret. Required. RUNNER_POOL_CREDENTIAL_FILE reads it from a mounted file instead (used only when the environment variable is empty). |
REELBOLT_WORKER_NAME | The name of this worker's enrolled row. Defaults to the hostname — so a VM whose hostname changed has silently lost its identity. |
REELBOLT_WORKER_SLOTS | Jobs run at once. Defaults to floor(vCPU / SANDBOX_CPU_LIMIT), never below 1 and never above 64. |
REELBOLT_WORKER_CPUS | Overrides the vCPU count, for a host or a container whose quota Go cannot see. |
REELBOLT_WORKER_SCRATCH | Per-job scratch root. Defaults to /var/lib/reelbolt/worker. |
REELBOLT_WORKER_DRAIN_SECONDS | How long the drain window is. Defaults to 120. |
SANDBOX_CONTAINER_CLI | docker (default), podman, or none for a machine with no container runtime. |
SANDBOX_IMAGE | The sandbox-runtime image. In the cloud this is the digest from sandbox/runtime-image.lock. |
SANDBOX_ROOT, SANDBOX_NETWORK, SANDBOX_CPU_LIMIT, SANDBOX_MEMORY_LIMIT, SANDBOX_PIDS_LIMIT, SANDBOX_TTL | The container driver's settings, read exactly as sandbox-executor reads them, because it is the same driver. |
DATABASE_URL | Read by reelbolt-worker enroll only. |
The vCPU count comes from REELBOLT_WORKER_CPUS, else the cgroup v2 cpu.max
quota, else the host's CPU count. The cgroup step is why this is not simply
runtime.NumCPU(): a worker shipped as a container with --cpus=4 on a
64-core VM must advertise 4 slots, not 32.
Capabilities: what a worker advertises, and why a digest matters
The capability document in the hello is what the gateway's claim query filters on, so the worker only ever builds an executor for a job type it actually advertised:
video.compileandvideo.analyze.extractneed ffmpeg. With a container runtime they run inside the runtime image, which is where the argv validator's guarantees are also enforced by the kernel: no network, a read-only root, the session directory as the only writable mount. Without one they run as host processes, which is what a desktop runner without Docker does.remotion.workspace-sessionadditionally needs a digest-pinned runtime image. A locally built tag has no digest, and the worker then advertises the type not at all — a render whose runtime is not pinned is a render this fleet cannot reproduce, and offering it would be a promise the worker cannot keep.
A worker that advertises no job types is not an error, and the gateway counts it
as connected-but-useless rather than as capacity. It does mean no work will ever
arrive, so the worker says so in the log at startup and reelbolt-worker caps
prints the same document without connecting.
Startup also creates the --internal sandbox network if it is missing (a pool VM
is not a compose stack, and a missing network would fail every Remotion job at
container-create time with an error that does not name it) and removes leftover
workspace containers from a previous process. If the network cannot be created,
the worker degrades to the ffmpeg job types and says why, rather than failing to
start.
Draining on SIGTERM
A spot reclaim or a scale-in sends SIGTERM. The worker then stops accepting work,
sends drain with its deadline and bye, closes, and exits — within
REELBOLT_WORKER_DRAIN_SECONDS (120 by default). Sessions it holds are not
waited for: the protocol's rule is that they are rehydrated elsewhere, so the
window bounds the drain, the goodbye and the close rather than a render.
Two consequences for the deployment: a systemd unit needs TimeoutStopSec
above the drain window, and a container platform needs a stop timeout above it
too. infra/worker/cloud-init.yaml sets 150 seconds.
The pool scaler: reading the scaling signal
reelbolt-scaler turns the gateway's answer into a node count:
RUNNER_GATEWAY_URL=http://runner-gateway:8090 \
RUNNER_INTERNAL_TOKEN=... \
POOL_SLOTS_PER_NODE=4 POOL_MAX_NODES=10 \
reelbolt-scaler -json
GET /internal/v1/scaling returns desiredCloudPoolSlots = queued + running + a
configurable buffer. The scaler reads that number and never recomputes it: a
second formula for the same quantity is a second thing to keep in step, and the
two would disagree the first time one of them changed. Its own arithmetic is only
the translation from slots to nodes, plus two rules:
- Never below what is running. Scaling a node away under a job that is
already running fails that job; it is not a saving.
runningNodesis a floor whateverPOOL_MAX_NODESsays. - A ceiling below that floor is reported, not obeyed. An operator whose cap is smaller than the fleet in use has a configuration problem, and terminating renders would hide it.
It also reads /metrics and cross-checks the fleet-wide
reelbolt_runner_queue_depth and reelbolt_runner_free_slots gauges against the
endpoint's per-replica counts. Those two are the gateway's registrations and the
scaler only ever parses them — the free-slot and connection counts in the
endpoint describe one replica, while the queue depth is database-wide, and a
scaler pointed at a single replica sees the whole queue and one pod's capacity.
That combination over-scales, and saying so out loud is cheaper than discovering
it on a bill.
Without -apply the scaler is a dry run. It prints the decision and changes
nothing. With -apply it runs POOL_SCALE_COMMAND through a shell with the
whole decision in its environment (POOL_NODES, POOL_DESIRED_SLOTS,
POOL_RUNNING, POOL_QUEUED, POOL_CAPPED_BY, ...), so a logged hook
invocation is a complete record of the round. Every value in that environment is
an integer or a fixed word this program computed.
What F12 must supply
There is no cloud credential in this repository and no provider API to call, so
this WP stops one step short of the provider on purpose. What exists is the whole
decision surface: the signal, the translation, the caps, the safety floor, the
warnings and a documented place to put the call. F12 must add exactly one
thing — a real provider action, i.e. POOL_SCALE_COMMAND replaced by a "set
the node count of this instance pool" call — plus the three things that call
implies and that no code in this repository can invent:
- The provider and its credential (an API token with instance-pool permissions, mounted as a secret).
- The cooldown and the convergence policy. This scaler is a single-shot decision, deliberately: it has no memory, so two runs three seconds apart can ask for two different node counts. A real autoscaler needs a damped or rate-limited control loop around it, and that is a property of how often the job runs, not of this binary.
- The cloud-init that makes a new node join the pool. A node that boots
without enrolling is refused 4401 —
infra/worker/cloud-init.yamlis the fragment F12's launch template must bake in, together with the pool credential and the pinned runtime-image digest.
POOL_MAX_NODES already exists in infra/compose/env.control.example (F14 put
it there), and the per-organisation CloudPool cap is F11's pool_slot_cap, which
the gateway enforces at claim time rather than the scaler.
Building the image and running the local smoke test
The image is the worker target in sandbox/Dockerfile and builds in the CI
docker matrix alongside the others (docs/ci.md). It carries the worker and
the scaler, the docker CLI and ffmpeg. It deliberately does not carry the
sandbox-runtime image: that is pulled by digest on the VM, so the render's
runtime is a pinned artefact rather than whatever a node happened to have. There
is no EXPOSE and no healthcheck, because a pool worker listens on nothing —
its liveness is the connection the gateway's own
reelbolt_runner_connected{kind="CloudPool"} gauge reports.
sandbox/scripts/worker-smoke.sh
stands up a real gateway, a throwaway Postgres and a real worker, and drives one
video.compile job end to end. It proves, with nothing faked:
- an unenrolled worker is refused 4401, and the refusal names the remedy;
enrollcreates exactly one row, and a second run lands on it;- the same binary then handshakes as a
CloudPoolrunner; - ffmpeg actually runs inside a real container of the sandbox-runtime image, encodes an MP4, and ffprobe reads its dimensions back through the session;
- the job completes as
Succeeded; - the scaling hook reads the endpoint and runs its command;
- SIGTERM drains the worker and it exits.
It does not cover remotion.workspace-session (needs a digest-pinned
runtime), presigned stage/publish (needs an object store with a public
endpoint), or any cloud provider. The first two are D8/D11 territory; the third
is F12's, and is the reason the previous section exists.
What it caught. Driving the real path is what found two defects that no fake
driver could: the media executor's default container CLI executed a binary named
run instead of the container CLI (every real media session failed while the
fake-CLI tests stayed green), and a session ended by the close op never
reached a terminal state — the job sat in Running until its lease expired and
the reaper called it Expired, so the engine would have been told
session_lost for a step that succeeded. Both are fixed and both now have
regression tests.