The compute fabric
The compute fabric is the part of ReelBolt that can run a workflow step somewhere other than the
engine process: on the customer's own paired desktop machine, or on a worker VM in our cloud pool.
Its control plane is the Go runner-gateway; its data plane is the runner protocol
(runner-protocol.md), which is what a runner speaks. The engine decides where
a step runs through one seam, IComputeTarget
(inference/src/ReelBolt.WorkflowEngine/Services/Compute/IComputeTarget.cs), and records the answer
on the step result. No executor resolves that seam yet — see "What is wired today and what is
not" before you conclude anything from the classes existing.
This page is the operator's and developer's map: what the three targets mean, what the selector decides and why, how to turn the fabric on, what is wired today and what is not, and the two deployment hazards that will bite an operator who skips the runbook.
Nothing here is required for a self-host install. Compute:Mode defaults to Local and the key
is deliberately absent from the shipped inference/src/ReelBolt.WorkflowEngine/appsettings.json, so
docker compose up needs no new configuration and the whole fabric is inert until an operator asks
for it. The reasoning is in the design record docs/design/runner-protocol-adr.md §1.
The three targets
ComputeTargetKind (inference/src/ReelBolt.Shared/Data/Models/ComputeTargetKind.cs) is what the
engine persists on workflow_step_results.compute_target. It is stored as an int and is
append-only: the numbers are the stored values.
| Target | Value | What it means | Trust |
|---|---|---|---|
Local | 0 | The engine ran the step in-process — the sandbox-executor plus in-process ffmpeg. This is what every install did before the fabric existed, and what Compute:Mode=Local produces. | This process |
CloudPool | 1 | One of our own worker VMs, serving any organisation's target=CloudPool jobs. | Trusted, but its output is verified anyway |
OrgRunner | 2 | The organisation's own paired machine (target=OrgRunner in the protocol). | Untrusted by construction — the binary is controlled by its owner — so its output is verified before promotion, tagged WorkflowStepResult.FromRunner and excluded from the step-result cache |
The protocol's two target strings ("CloudPool" and "OrgRunner") are the same three-way choice
minus Local, which has no protocol representation at all — a local step is not a job. The mapping
between the two vocabularies lives in one place, ComputeTargets in
inference/src/ReelBolt.WorkflowEngine/Services/Compute/IComputeTarget.cs.
How to turn the fabric on
Two settings, and one of them is on the gateway's side:
Compute:Mode=Fabricon the WorkflowEngine. That is the switch. With the key absent (the shipped state) the selector answersLocalbefore it reads the database or calls the gateway, so an install that has never heard of the fabric behaves exactly as it did before.Compute:InternalToken, or theRUNNER_INTERNAL_TOKENenvironment variable — the same variable the gateway reads, so one value configures both sides. It is never logged. A gateway with no token configured answers503 internal_api_disabledon its whole/internal/v1surface; an engine with no token refuses every call rather than sending one unauthenticated.
Optionally Compute:GatewayBaseUrl (default http://runner-gateway:8090, the compose service name)
and the tuning keys RunnerQueryCacheSeconds (5), RequestTimeoutSeconds (30), JobWaitSeconds
(30) and SessionWaitSeconds (900).
The gateway needs DATABASE_URL, and it will not start without it: it authenticates every desktop
runner against runner_devices.public_key and re-checks revoked_at on live connections. It never
migrates — the Inference API owns runner_devices and platform_settings, the engine owns
runner_jobs and runner_sessions.
Turning Mode on does not, today, route any step to a runner. See the next section.
What is wired today and what is not
This is the part a future reader must not have to rediscover, so it is stated explicitly rather than implied by which classes exist.
Wired and verified. The selector, the gateway client, the per-execution session gate and the
output verifier are all registered in the engine's composition root
(WorkflowEngineServiceCollectionExtensions), and ComputeFabricCompositionTests proves they
resolve from the real AddWorkflowEngineServices — including that the engine comes up in Local
mode with no Compute configuration at all. The gateway itself (job queue, leases, fencing,
routing, the scaling endpoint, the kill switches) is complete and exercised against a real broker,
a real PostgreSQL and real WebSocket clients.
Not wired. The callers landed after this document was written — D7b routes the Remotion
workspace tools through RemoteSandboxSession and its journal, D11 routes the media tools through
RemoteMediaWorkspace, D8 publishes renders, voiceover WAVs and project sources through presigned
URLs into the job's staging key space, and D11's remote publish is one of the two places that calls
D12's IOutputVerifier (D8's render path is the other). They are all reached from Execution/, not
only from the DI registrations, so setting Compute:Mode=Fabric now routes real work.
Consequences worth stating plainly:
- A
CloudPooljob created at the gateway staysQueuedforever: see "Nothing enrolls CloudPool workers" below. That is a missing data plane (no worker runs in this tree), not a missing caller. video.analyze.extractnow has its engine caller.VideoAnalyze's extraction phase names every file it stages, writes, reads and publishes by a session-relative virtual path and reaches ffmpeg only throughIMediaWorkspace, so a Fabric deployment sends the decode, the shot and silence detection, the WAV, the frame grids, the chroma tracking and the caption keyframes to the runner. The inference phase (ASR, vision captioning, subject location, the derush pre-pass, input screening) stays on the enginebecause it needs provider credentials— asrunner-protocol.mdstates — and reads the small media it consumes (each caption keyframe, the extracted WAV) back through the session's ownreadop, so the footage itself never crosses back. The ONE configuration that still runs locally isTrackObjects: markerless object tracking samples up to 600 MB of grayscale frames across a whole clip and so needs the source on the machine running it, which is the same "platform-only in v1" decisionrunner-protocol.mdrecords;RunnerEligibleis false for it and the selector answers Local rather than failing the step. Subject location is not in that set: it consumes one small keyframe JPEG per shot, which the workspace carries as bytes.
The selector's decision table
ComputeTargetSelector (inference/src/ReelBolt.WorkflowEngine/Services/Compute/ComputeTargetSelector.cs)
answers one question — "where does this unit of work run?" — from four inputs: the deployment mode,
the organisation's compute_preference, the plan entitlement and whether an eligible runner is
online right now. It returns a ComputeDecision, a value: nothing is dispatched by producing one.
With Compute:Mode=Local (the default) it answers Local immediately, before any database read
and before any gateway call.
With Fabric:
| Condition | Answer |
|---|---|
The step is not runner-eligible in v1 (a VideoAnalyze with TrackObjects) | Local — it runs where it always has. It does not fail |
| The org may use neither a runner nor the pool | COMPUTE_NOT_AVAILABLE (a coded step failure) |
Preference CloudOnly (and any preference downgraded to it because the plan withholds runner use) | CloudPool when the pool is allowed, otherwise COMPUTE_NOT_AVAILABLE |
Preference RunnerOnly, an eligible desktop runner online | OrgRunner |
Preference RunnerOnly, no eligible runner | NO_RUNNER_ONLINE — the step FAILS. There is no silent cloud fallback; that is the whole point of the preference |
Preference PreferRunner, runner online | OrgRunner |
Preference PreferRunner, no runner, pool allowed | CloudPool |
Preference Auto with a runner online | OrgRunner (Auto behaves as PreferRunner) |
Preference Auto with no runner, pool allowed | CloudPool (Auto behaves as CloudOnly until a runner is paired) |
PreferRunner/Auto, no runner, pool not allowed | NO_RUNNER_ONLINE |
503 and transport failure are "the fabric is not deployed", and never fail a step
The gateway client maps DNS failure, connection refused, TLS failure, a client timeout and the
gateway's 503 internal_api_disabled onto one exception, ComputeFabricUnavailableException — a
deployment fact, deliberately not an HttpRequestException that would turn "the fabric is not
deployed here" into a failed workflow. The selector degrades on it:
- Under any preference except
RunnerOnly, the step runsLocal. A gateway we cannot reach cannot take aCloudPooljob either, soLocalis the only answer that actually runs the work. - Under
RunnerOnlyit still reportsNO_RUNNER_ONLINE. A workspace that asked for its own machine or nothing gets the coded failure, not a silent downgrade to platform compute.
An engine with Fabric set but no internal token is classified the same way, deliberately: the
caller's recovery is identical.
A 4xx refusal (a wrong token, a malformed query, a body the client cannot parse) is a different
class — RunnerGatewayRequestException. The selector logs it as a warning and reads it as "no
runner confirmed", never as a step failure: a refusal says nothing about whether a runner exists.
The runner-presence answer is cached for Compute:RunnerQueryCacheSeconds (5 s), including a
negative and an unavailable one, so a wrong token cannot put a gateway call on the path of every
step. Over one execution that means an outage can be noticed up to five seconds late.
The pool-worker trap the selector avoids itself
GET /internal/v1/runners?org= returns CloudPool workers too: a pool worker's orgId is null and
the gateway's filter only excludes entries whose org is non-empty and different. Counting one as
"this org has a runner" would send the org's work to our own pool while reporting it as local
compute. The selector therefore checks Kind == Desktop itself rather than trusting the query
filter, and a test pins it.
The per-execution session gate
runner_sessions carries a partial UNIQUE(execution_id) WHERE closed_at IS NULL, so at most one
session row may be open for a workflow execution at a time. The gateway opens that row when a
runner is assigned a job and deliberately tolerates a violation (a savepoint, logged, and the job
still runs — just without a session row of its own), because the constraint belongs to the schema.
That makes the invariant the engine's to keep, and it matters for a concrete reason: a job with
no session row has no runner_sessions.gateway_instance for D5c's op routing to use, so the engine
would have no way to address the runner executing it.
IExecutionSessionGate (inference/src/ReelBolt.WorkflowEngine/Services/Compute/ExecutionSessionGate.cs)
is that rule, expressed once: it hands out at most one lease per execution id, held from the
moment the job is created (a runner may be assigned it at any instant) until the job reaches a
terminal state. Different executions never queue behind each other, so one busy tenant cannot starve
another.
What it costs. The two halves cannot be separated: RunnerGatewayClient.OpenJobAsync takes the
execution's slot and creates the job in the same call, so a second session of the same execution
queues behind the first rather than failing fast. The queue is bounded by
Compute:SessionWaitSeconds (900 s by default); when it expires the caller gets
ComputeSessionBusyException and the step fails with RUNNER_SESSION_BUSY. A workflow that
genuinely needed two concurrent runner sessions in one execution would wait for itself until that
bound, then fail — which is the intended outcome, because the schema forbids the alternative.
The guarantee is in-process, and that is sufficient: an execution is run by exactly one engine
instance (the WorkflowEngine:Lease execution lease establishes it), so a second instance cannot be
running the same execution's steps concurrently. The gateway's own tolerance stays the backstop for
everything the gate cannot see — a job created by an older build, a run migrated mid-flight.
Output verification
Runner output is untrusted, and verification happens before promotion, not after:
- The engine computes the object's SHA-256 itself; a runner-reported digest is treated as a hint, never as the answer.
- The object is probed with
ffprobeover a presigned GET of the staging object using an HTTP range, so nothing is downloaded: duration (± one frame, else 0.25 s), width, height, codec and frame rate are compared against the platform-side job spec. - A spot check (
Compute:OutputVerification:SpotCheckPercent, default 10) runsblackdetectandfreezedetecton three random one-second windows — onOrgRunneroutput only. ACloudPoolworker is ours, so its output is trusted but still verified. - A failure deletes the staged object and returns a
RUNNER_OUTPUT_*code rather than throwing.
Two billing consequences ride on the same fact. FromRunner is derived from
ComputeTarget == OrgRunner, never set independently, so the tag and the target cannot disagree;
and every RUNNER_OUTPUT_* code is classified FailureClass.Runner, which
CreditGate.IsRefundable explicitly maps to not refundable — unmapped failure codes refund by
default, which would be the wrong answer for a runner's bad output.
The self-host matrix
| Deployment | Compute:Mode | Gateway | Runner | What runs the work |
|---|---|---|---|---|
docker compose up (default, and the only supported self-host path) | absent → Local | not needed, not started | none | sandbox-executor + in-process ffmpeg |
| Self-host with a paired desktop machine (not reachable today) | Fabric | required | the paired device | OrgRunner |
| Our cloud | Fabric | required | cloud workers | CloudPool, OrgRunner when a customer has one |
Self-host keeps the Local target. Retiring sandbox-executor in favour of an embedded worker
is a post-launch follow-up, not v1, and no part of this document asks a self-hoster to run a gateway.
The optional fabric compose profile exists for developers only, and the default
docker compose up does not start it — see "Running the gateway locally".
Deployment hazards
x-max-priority is a queue ARGUMENT, and an existing broker can refuse it
The engine declares the workflow-execution queue with EnablePriority(...) so that plan priority
(B8c) actually reorders the queue. x-max-priority is a queue argument, and a broker answers a
re-declaration that differs from the existing queue with PRECONDITION_FAILED (406).
The consequence is a real deployment hazard: an engine upgrading onto a broker that already holds
the argument-less workflow-execution queue will fail to start its bus until that queue is deleted.
Drain it first — messages sitting in the old queue are not migrated.
Escape hatch: WorkflowEngine:QueuePriority:MaxPriority=0 (compose:
WORKFLOW_QUEUE_MAX_PRIORITY=0) keeps the pre-B8c declaration and disables the ordering. It is the
only value that does.
Nothing enrolls CloudPool workers, so the pool is unreachable
The gateway refuses a pool worker with 4401 unless a runner_devices row exists for it
(kind='CloudPool', name = hello.auth.workerName) — fail-closed rather than inventing a
synthetic identity. No code in this tree creates that row. D5a flagged it as a hard
prerequisite and it is still true: D13 (with F) is what enrolls pool rows, and sandbox/cmd/reelbolt-worker
does not exist yet.
So a pool job created through the internal API is created successfully and then stays Queued.
Nothing errors — which is exactly why it is written down here. Pairing a desktop runner is a
different path and does write a runner_devices row.
Other things that make a fabric look broken when it is not
RUNNER_POOL_CREDENTIALSempty: logged at startup, and no CloudPool worker can authenticate.RUNNER_JOB_TOKEN_KEYunset: fine for one replica; with several, a runner that reconnects to a different replica is told its lease was lost. Set it for any multi-replica deployment.RABBITMQ_URLunset: theRunnerDeviceRevokedfast path is off and revocation relies on the 60-secondrevoked_atre-check, which is the protocol's own authoritative backstop.- Kill switches: an unreadable
platform_settingsis PAUSED everywhere, in the engine and on the gateway, and a disabled organisation makes the gateway refuse a hello with4403— which is terminal for the runner, so the gateway applies it only to a value it actually read. See platform-kill-switches.md.
Running the gateway locally
The gateway image is the runner-gateway target of sandbox/Dockerfile: an Alpine image holding
one WebSocket endpoint plus /healthz and /metrics, with no Docker CLI, no Node and no sandbox
tooling. It is the only image in this tree that faces the public internet, so what it does not
contain matters as much as what it does.
An optional compose profile starts it for a developer:
# Default stack is unchanged: no gateway, Compute:Mode=Local.
docker compose up -d
# Start the gateway too. Its service name is what Compute:GatewayBaseUrl already names.
docker compose --profile fabric up -d runner-gateway
The profile contains only the gateway. The worker half of the plan's profile is D13 and there is
no reelbolt-worker binary to start — see the hazard above. Starting the gateway alone is still
useful: it is the engine's default Compute:GatewayBaseUrl, it exposes the F11 switch metrics, and
it is where the 503 internal_api_disabled shape can be observed directly.
To make the engine actually call it, add COMPUTE_MODE=Fabric and a matching
RUNNER_INTERNAL_TOKEN to .env. Both are passed through to both services, so one value configures
both sides. The engine still routes nothing to a runner — see "What is wired today and what is
not".
For the public path, nginx already ships nginx/runner-server.conf.template: it is rendered only
when UPSTREAM_RUNNER_GATEWAY is set, so with no runner deployment there is no wss://runner.<domain>
server block at all.
Verifying a deployment
# The gateway is up and holds no runner.
curl -s localhost:8090/healthz
# Its internal API is open, and says which runners are online (needs the shared token).
curl -s -H "Authorization: Bearer $RUNNER_INTERNAL_TOKEN" \
'http://localhost:8090/internal/v1/runners?org=<org-guid>'
# -> 401 without the bearer; 503 {"error":"internal_api_disabled"} when the GATEWAY has no token.
# The kill-switch readers agree, and neither is guessing.
curl -s localhost:8090/metrics | grep reelbolt_platform_switches_unreadable_total
curl -s localhost:8090/metrics | grep reelbolt_runner_connected
# Which target each step actually ran on.
psql "$DATABASE_URL" -c "select compute_target, from_runner, count(*) from workflow_step_results group by 1,2;"
Published in a real run of this image against a throwaway PostgreSQL: /healthz is 200, the
internal API answers 503 internal_api_disabled with no RUNNER_INTERNAL_TOKEN and 401 without a
bearer when it has one, GET /internal/v1/runners answers {"online":0,"runners":[]}, and two
POST /internal/v1/jobs with the same idempotency key produce exactly one runner_jobs row
(201 then 200, the same jobId). Dropping platform_settings at runtime logs
platform switches are unreadable, increments reelbolt_platform_switches_unreadable_total, and
leaves the gateway serving while it assigns nothing — the "unreadable means paused, but never a
refused hello" rule, observed rather than described.
Failure codes a compute decision produces
These reach the user the same way every other engine failure code does, in the step's error details. The classification is asserted by the B6b drift guard, so adding one without classifying it fails the build.
| Code | When | Failure class |
|---|---|---|
NO_RUNNER_ONLINE | RunnerOnly (or an unentitled PreferRunner/Auto) with no eligible runner online, including when the fabric itself is unreachable | UserInput — the user's own choice coming back at them |
COMPUTE_NOT_AVAILABLE | The organisation may use neither a runner nor the pool | EntitlementDenied |
RUNNER_SESSION_BUSY | The Compute:SessionWaitSeconds wait for the execution's single session slot expired | PlatformInternal |
RUNNER_OUTPUT_* | Output verification failed: missing, too large, unreadable, duration/frame-size/codec/frame-rate mismatch, checksum mismatch, blank, promote failed | Runner (except RUNNER_OUTPUT_PROMOTE_FAILED, which is PlatformInternal) |
None of these is refunded as a platform fault: see metering-and-billing.md.