Skip to main content

The compute fabric

The compute fabric is the part of ReelBolt that can run a workflow step somewhere other than the engine process: on the customer's own paired desktop machine, or on a worker VM in our cloud pool. Its control plane is the Go runner-gateway; its data plane is the runner protocol (runner-protocol.md), which is what a runner speaks. The engine decides where a step runs through one seam, IComputeTarget (inference/src/ReelBolt.WorkflowEngine/Services/Compute/IComputeTarget.cs), and records the answer on the step result. No executor resolves that seam yet — see "What is wired today and what is not" before you conclude anything from the classes existing.

This page is the operator's and developer's map: what the three targets mean, what the selector decides and why, how to turn the fabric on, what is wired today and what is not, and the two deployment hazards that will bite an operator who skips the runbook.

Nothing here is required for a self-host install. Compute:Mode defaults to Local and the key is deliberately absent from the shipped inference/src/ReelBolt.WorkflowEngine/appsettings.json, so docker compose up needs no new configuration and the whole fabric is inert until an operator asks for it. The reasoning is in the design record docs/design/runner-protocol-adr.md §1.

The three targets​

ComputeTargetKind (inference/src/ReelBolt.Shared/Data/Models/ComputeTargetKind.cs) is what the engine persists on workflow_step_results.compute_target. It is stored as an int and is append-only: the numbers are the stored values.

TargetValueWhat it meansTrust
Local0The engine ran the step in-process — the sandbox-executor plus in-process ffmpeg. This is what every install did before the fabric existed, and what Compute:Mode=Local produces.This process
CloudPool1One of our own worker VMs, serving any organisation's target=CloudPool jobs.Trusted, but its output is verified anyway
OrgRunner2The organisation's own paired machine (target=OrgRunner in the protocol).Untrusted by construction — the binary is controlled by its owner — so its output is verified before promotion, tagged WorkflowStepResult.FromRunner and excluded from the step-result cache

The protocol's two target strings ("CloudPool" and "OrgRunner") are the same three-way choice minus Local, which has no protocol representation at all — a local step is not a job. The mapping between the two vocabularies lives in one place, ComputeTargets in inference/src/ReelBolt.WorkflowEngine/Services/Compute/IComputeTarget.cs.

How to turn the fabric on​

Two settings, and one of them is on the gateway's side:

  1. Compute:Mode=Fabric on the WorkflowEngine. That is the switch. With the key absent (the shipped state) the selector answers Local before it reads the database or calls the gateway, so an install that has never heard of the fabric behaves exactly as it did before.
  2. Compute:InternalToken, or the RUNNER_INTERNAL_TOKEN environment variable — the same variable the gateway reads, so one value configures both sides. It is never logged. A gateway with no token configured answers 503 internal_api_disabled on its whole /internal/v1 surface; an engine with no token refuses every call rather than sending one unauthenticated.

Optionally Compute:GatewayBaseUrl (default http://runner-gateway:8090, the compose service name) and the tuning keys RunnerQueryCacheSeconds (5), RequestTimeoutSeconds (30), JobWaitSeconds (30) and SessionWaitSeconds (900).

The gateway needs DATABASE_URL, and it will not start without it: it authenticates every desktop runner against runner_devices.public_key and re-checks revoked_at on live connections. It never migrates — the Inference API owns runner_devices and platform_settings, the engine owns runner_jobs and runner_sessions.

Turning Mode on does not, today, route any step to a runner. See the next section.

What is wired today and what is not​

This is the part a future reader must not have to rediscover, so it is stated explicitly rather than implied by which classes exist.

Wired and verified. The selector, the gateway client, the per-execution session gate and the output verifier are all registered in the engine's composition root (WorkflowEngineServiceCollectionExtensions), and ComputeFabricCompositionTests proves they resolve from the real AddWorkflowEngineServices — including that the engine comes up in Local mode with no Compute configuration at all. The gateway itself (job queue, leases, fencing, routing, the scaling endpoint, the kill switches) is complete and exercised against a real broker, a real PostgreSQL and real WebSocket clients.

Not wired. The callers landed after this document was written — D7b routes the Remotion workspace tools through RemoteSandboxSession and its journal, D11 routes the media tools through RemoteMediaWorkspace, D8 publishes renders, voiceover WAVs and project sources through presigned URLs into the job's staging key space, and D11's remote publish is one of the two places that calls D12's IOutputVerifier (D8's render path is the other). They are all reached from Execution/, not only from the DI registrations, so setting Compute:Mode=Fabric now routes real work.

Consequences worth stating plainly:

  • A CloudPool job created at the gateway stays Queued forever: see "Nothing enrolls CloudPool workers" below. That is a missing data plane (no worker runs in this tree), not a missing caller.
  • video.analyze.extract now has its engine caller. VideoAnalyze's extraction phase names every file it stages, writes, reads and publishes by a session-relative virtual path and reaches ffmpeg only through IMediaWorkspace, so a Fabric deployment sends the decode, the shot and silence detection, the WAV, the frame grids, the chroma tracking and the caption keyframes to the runner. The inference phase (ASR, vision captioning, subject location, the derush pre-pass, input screening) stays on the engine because it needs provider credentials — as runner-protocol.md states — and reads the small media it consumes (each caption keyframe, the extracted WAV) back through the session's own read op, so the footage itself never crosses back. The ONE configuration that still runs locally is TrackObjects: markerless object tracking samples up to 600 MB of grayscale frames across a whole clip and so needs the source on the machine running it, which is the same "platform-only in v1" decision runner-protocol.md records; RunnerEligible is false for it and the selector answers Local rather than failing the step. Subject location is not in that set: it consumes one small keyframe JPEG per shot, which the workspace carries as bytes.

The selector's decision table​

ComputeTargetSelector (inference/src/ReelBolt.WorkflowEngine/Services/Compute/ComputeTargetSelector.cs) answers one question — "where does this unit of work run?" — from four inputs: the deployment mode, the organisation's compute_preference, the plan entitlement and whether an eligible runner is online right now. It returns a ComputeDecision, a value: nothing is dispatched by producing one.

With Compute:Mode=Local (the default) it answers Local immediately, before any database read and before any gateway call.

With Fabric:

ConditionAnswer
The step is not runner-eligible in v1 (a VideoAnalyze with TrackObjects)Local — it runs where it always has. It does not fail
The org may use neither a runner nor the poolCOMPUTE_NOT_AVAILABLE (a coded step failure)
Preference CloudOnly (and any preference downgraded to it because the plan withholds runner use)CloudPool when the pool is allowed, otherwise COMPUTE_NOT_AVAILABLE
Preference RunnerOnly, an eligible desktop runner onlineOrgRunner
Preference RunnerOnly, no eligible runnerNO_RUNNER_ONLINE — the step FAILS. There is no silent cloud fallback; that is the whole point of the preference
Preference PreferRunner, runner onlineOrgRunner
Preference PreferRunner, no runner, pool allowedCloudPool
Preference Auto with a runner onlineOrgRunner (Auto behaves as PreferRunner)
Preference Auto with no runner, pool allowedCloudPool (Auto behaves as CloudOnly until a runner is paired)
PreferRunner/Auto, no runner, pool not allowedNO_RUNNER_ONLINE

503 and transport failure are "the fabric is not deployed", and never fail a step​

The gateway client maps DNS failure, connection refused, TLS failure, a client timeout and the gateway's 503 internal_api_disabled onto one exception, ComputeFabricUnavailableException — a deployment fact, deliberately not an HttpRequestException that would turn "the fabric is not deployed here" into a failed workflow. The selector degrades on it:

  • Under any preference except RunnerOnly, the step runs Local. A gateway we cannot reach cannot take a CloudPool job either, so Local is the only answer that actually runs the work.
  • Under RunnerOnly it still reports NO_RUNNER_ONLINE. A workspace that asked for its own machine or nothing gets the coded failure, not a silent downgrade to platform compute.

An engine with Fabric set but no internal token is classified the same way, deliberately: the caller's recovery is identical.

A 4xx refusal (a wrong token, a malformed query, a body the client cannot parse) is a different class — RunnerGatewayRequestException. The selector logs it as a warning and reads it as "no runner confirmed", never as a step failure: a refusal says nothing about whether a runner exists.

The runner-presence answer is cached for Compute:RunnerQueryCacheSeconds (5 s), including a negative and an unavailable one, so a wrong token cannot put a gateway call on the path of every step. Over one execution that means an outage can be noticed up to five seconds late.

The pool-worker trap the selector avoids itself​

GET /internal/v1/runners?org= returns CloudPool workers too: a pool worker's orgId is null and the gateway's filter only excludes entries whose org is non-empty and different. Counting one as "this org has a runner" would send the org's work to our own pool while reporting it as local compute. The selector therefore checks Kind == Desktop itself rather than trusting the query filter, and a test pins it.

The per-execution session gate​

runner_sessions carries a partial UNIQUE(execution_id) WHERE closed_at IS NULL, so at most one session row may be open for a workflow execution at a time. The gateway opens that row when a runner is assigned a job and deliberately tolerates a violation (a savepoint, logged, and the job still runs — just without a session row of its own), because the constraint belongs to the schema.

That makes the invariant the engine's to keep, and it matters for a concrete reason: a job with no session row has no runner_sessions.gateway_instance for D5c's op routing to use, so the engine would have no way to address the runner executing it.

IExecutionSessionGate (inference/src/ReelBolt.WorkflowEngine/Services/Compute/ExecutionSessionGate.cs) is that rule, expressed once: it hands out at most one lease per execution id, held from the moment the job is created (a runner may be assigned it at any instant) until the job reaches a terminal state. Different executions never queue behind each other, so one busy tenant cannot starve another.

What it costs. The two halves cannot be separated: RunnerGatewayClient.OpenJobAsync takes the execution's slot and creates the job in the same call, so a second session of the same execution queues behind the first rather than failing fast. The queue is bounded by Compute:SessionWaitSeconds (900 s by default); when it expires the caller gets ComputeSessionBusyException and the step fails with RUNNER_SESSION_BUSY. A workflow that genuinely needed two concurrent runner sessions in one execution would wait for itself until that bound, then fail — which is the intended outcome, because the schema forbids the alternative.

The guarantee is in-process, and that is sufficient: an execution is run by exactly one engine instance (the WorkflowEngine:Lease execution lease establishes it), so a second instance cannot be running the same execution's steps concurrently. The gateway's own tolerance stays the backstop for everything the gate cannot see — a job created by an older build, a run migrated mid-flight.

Output verification​

Runner output is untrusted, and verification happens before promotion, not after:

  • The engine computes the object's SHA-256 itself; a runner-reported digest is treated as a hint, never as the answer.
  • The object is probed with ffprobe over a presigned GET of the staging object using an HTTP range, so nothing is downloaded: duration (± one frame, else 0.25 s), width, height, codec and frame rate are compared against the platform-side job spec.
  • A spot check (Compute:OutputVerification:SpotCheckPercent, default 10) runs blackdetect and freezedetect on three random one-second windows — on OrgRunner output only. A CloudPool worker is ours, so its output is trusted but still verified.
  • A failure deletes the staged object and returns a RUNNER_OUTPUT_* code rather than throwing.

Two billing consequences ride on the same fact. FromRunner is derived from ComputeTarget == OrgRunner, never set independently, so the tag and the target cannot disagree; and every RUNNER_OUTPUT_* code is classified FailureClass.Runner, which CreditGate.IsRefundable explicitly maps to not refundable — unmapped failure codes refund by default, which would be the wrong answer for a runner's bad output.

The self-host matrix​

DeploymentCompute:ModeGatewayRunnerWhat runs the work
docker compose up (default, and the only supported self-host path)absent → Localnot needed, not startednonesandbox-executor + in-process ffmpeg
Self-host with a paired desktop machine (not reachable today)Fabricrequiredthe paired deviceOrgRunner
Our cloudFabricrequiredcloud workersCloudPool, OrgRunner when a customer has one

Self-host keeps the Local target. Retiring sandbox-executor in favour of an embedded worker is a post-launch follow-up, not v1, and no part of this document asks a self-hoster to run a gateway.

The optional fabric compose profile exists for developers only, and the default docker compose up does not start it — see "Running the gateway locally".

Deployment hazards​

x-max-priority is a queue ARGUMENT, and an existing broker can refuse it​

The engine declares the workflow-execution queue with EnablePriority(...) so that plan priority (B8c) actually reorders the queue. x-max-priority is a queue argument, and a broker answers a re-declaration that differs from the existing queue with PRECONDITION_FAILED (406).

The consequence is a real deployment hazard: an engine upgrading onto a broker that already holds the argument-less workflow-execution queue will fail to start its bus until that queue is deleted. Drain it first — messages sitting in the old queue are not migrated.

Escape hatch: WorkflowEngine:QueuePriority:MaxPriority=0 (compose: WORKFLOW_QUEUE_MAX_PRIORITY=0) keeps the pre-B8c declaration and disables the ordering. It is the only value that does.

Nothing enrolls CloudPool workers, so the pool is unreachable​

The gateway refuses a pool worker with 4401 unless a runner_devices row exists for it (kind='CloudPool', name = hello.auth.workerName) — fail-closed rather than inventing a synthetic identity. No code in this tree creates that row. D5a flagged it as a hard prerequisite and it is still true: D13 (with F) is what enrolls pool rows, and sandbox/cmd/reelbolt-worker does not exist yet.

So a pool job created through the internal API is created successfully and then stays Queued. Nothing errors — which is exactly why it is written down here. Pairing a desktop runner is a different path and does write a runner_devices row.

Other things that make a fabric look broken when it is not​

  • RUNNER_POOL_CREDENTIALS empty: logged at startup, and no CloudPool worker can authenticate.
  • RUNNER_JOB_TOKEN_KEY unset: fine for one replica; with several, a runner that reconnects to a different replica is told its lease was lost. Set it for any multi-replica deployment.
  • RABBITMQ_URL unset: the RunnerDeviceRevoked fast path is off and revocation relies on the 60-second revoked_at re-check, which is the protocol's own authoritative backstop.
  • Kill switches: an unreadable platform_settings is PAUSED everywhere, in the engine and on the gateway, and a disabled organisation makes the gateway refuse a hello with 4403 — which is terminal for the runner, so the gateway applies it only to a value it actually read. See platform-kill-switches.md.

Running the gateway locally​

The gateway image is the runner-gateway target of sandbox/Dockerfile: an Alpine image holding one WebSocket endpoint plus /healthz and /metrics, with no Docker CLI, no Node and no sandbox tooling. It is the only image in this tree that faces the public internet, so what it does not contain matters as much as what it does.

An optional compose profile starts it for a developer:

# Default stack is unchanged: no gateway, Compute:Mode=Local.
docker compose up -d

# Start the gateway too. Its service name is what Compute:GatewayBaseUrl already names.
docker compose --profile fabric up -d runner-gateway

The profile contains only the gateway. The worker half of the plan's profile is D13 and there is no reelbolt-worker binary to start — see the hazard above. Starting the gateway alone is still useful: it is the engine's default Compute:GatewayBaseUrl, it exposes the F11 switch metrics, and it is where the 503 internal_api_disabled shape can be observed directly.

To make the engine actually call it, add COMPUTE_MODE=Fabric and a matching RUNNER_INTERNAL_TOKEN to .env. Both are passed through to both services, so one value configures both sides. The engine still routes nothing to a runner — see "What is wired today and what is not".

For the public path, nginx already ships nginx/runner-server.conf.template: it is rendered only when UPSTREAM_RUNNER_GATEWAY is set, so with no runner deployment there is no wss://runner.<domain> server block at all.

Verifying a deployment​

# The gateway is up and holds no runner.
curl -s localhost:8090/healthz

# Its internal API is open, and says which runners are online (needs the shared token).
curl -s -H "Authorization: Bearer $RUNNER_INTERNAL_TOKEN" \
'http://localhost:8090/internal/v1/runners?org=<org-guid>'
# -> 401 without the bearer; 503 {"error":"internal_api_disabled"} when the GATEWAY has no token.

# The kill-switch readers agree, and neither is guessing.
curl -s localhost:8090/metrics | grep reelbolt_platform_switches_unreadable_total
curl -s localhost:8090/metrics | grep reelbolt_runner_connected

# Which target each step actually ran on.
psql "$DATABASE_URL" -c "select compute_target, from_runner, count(*) from workflow_step_results group by 1,2;"

Published in a real run of this image against a throwaway PostgreSQL: /healthz is 200, the internal API answers 503 internal_api_disabled with no RUNNER_INTERNAL_TOKEN and 401 without a bearer when it has one, GET /internal/v1/runners answers {"online":0,"runners":[]}, and two POST /internal/v1/jobs with the same idempotency key produce exactly one runner_jobs row (201 then 200, the same jobId). Dropping platform_settings at runtime logs platform switches are unreadable, increments reelbolt_platform_switches_unreadable_total, and leaves the gateway serving while it assigns nothing — the "unreadable means paused, but never a refused hello" rule, observed rather than described.

Failure codes a compute decision produces​

These reach the user the same way every other engine failure code does, in the step's error details. The classification is asserted by the B6b drift guard, so adding one without classifying it fails the build.

CodeWhenFailure class
NO_RUNNER_ONLINERunnerOnly (or an unentitled PreferRunner/Auto) with no eligible runner online, including when the fabric itself is unreachableUserInput — the user's own choice coming back at them
COMPUTE_NOT_AVAILABLEThe organisation may use neither a runner nor the poolEntitlementDenied
RUNNER_SESSION_BUSYThe Compute:SessionWaitSeconds wait for the execution's single session slot expiredPlatformInternal
RUNNER_OUTPUT_*Output verification failed: missing, too large, unreadable, duration/frame-size/codec/frame-rate mismatch, checksum mismatch, blank, promote failedRunner (except RUNNER_OUTPUT_PROMOTE_FAILED, which is PlatformInternal)

None of these is refunded as a platform fault: see metering-and-billing.md.