Continuous Integration
CI builds and tests everything; a push to master also publishes signed images to
GHCR, and a separate workflow deploys them to the cloud VMs and the marketing site to
Cloudflare Pages (see "Deployment" below). Nothing deploys until the founder has created
the GitHub environments and secrets: every deploy job skips with a notice when they are
absent, so CI stays green before the accounts exist.
| Workflow | File | Trigger |
|---|---|---|
| CI | .github/workflows/ci.yml | push to master, every PR, manual |
| Deploy | .github/workflows/deploy.yml | green CI on master (staging), manual dispatch (production) |
| Runtime image release | .github/workflows/runtime-image-release.yml | runtime-v* tag (the desktop runner's sandbox runtime image) |
| Runner release | .github/workflows/runner-release.yml | runner-v* tag (publishes the desktop runner to the beta channel), manual dispatch (promote to stable, or a build-and-verify dry run) |
There is no CodeQL workflow. Code scanning on a private repository requires GitHub Code Security, which this repo does not have, so CodeQL could analyse but never upload and failed every run; it was removed rather than left red.
The runner release is the one workflow that does not follow the "skip with a
notice until the accounts exist" rule the deploy jobs follow. It runs when
somebody asks for a release, not on every push, so a release that silently does
nothing is worse than one that names what is missing: its publish job fails
with the checklist rather than skipping. The build legs still run, so the
toolchain and the packaging are exercisable before the release key, the Windows
certificate and the R2 credentials exist. See
runner/build/README.md.
What runs
| Job | Covers | Checks |
|---|---|---|
go (matrix: api, sandbox, runner) | the three Go modules | go mod tidy is a no-op, gofmt, go vet, go build ./..., go test ./... -race + coverage artifact |
runner-windows | runner/ | go vet and go test -race on windows-latest (runtime detection is OS-specific). The go job's runner leg also cross-compiles for windows/amd64 and linux/arm64 |
golangci-lint (matrix: api, sandbox, runner) | the three Go modules | golangci-lint v2.13.2, config in .golangci.yml |
dotnet | inference/ReelBolt.sln (Shared, Inference.Api, WorkflowEngine, WorkflowEngine.Tests) | dotnet restore/build -c Release (Roslyn + .NET analyzers run here), dotnet test with TRX + Cobertura coverage artifacts |
dotnet-format | same solution | dotnet format --verify-no-changes --severity warn — blocking, and one of the jobs the ci aggregate requires |
node (matrix: web, site) | both Next.js apps | npm ci, tsc --noEmit, npm run lint (ESLint), npm test (Vitest, web only), npm run build |
docs-site | docs/ + docs-site/ | npm ci, npm run build (Docusaurus — fails on any broken link, so it doubles as the docs link check) |
docker (matrix: 12 images, incl. mcp) | every Dockerfile | buildx build with a GHA layer cache per image; on push to master only it also pushes ghcr.io/vecchiotom/reelbolt-<name>:<sha> and :master with an SBOM and provenance attestation and signs the digest with cosign (keyless, GitHub OIDC; packages: write/id-token: write are scoped to this job, PRs never push). Builds web/site/docs-site at their production target and sandbox at sandbox-runtime, control-plane, runner-gateway and worker (the D13 CloudPool worker, whose image is the only CI-built one that runs customer renders); inference-api and docs-site get the /docs tree through a named build-contexts entry |
sandbox-runtime-arm64 | sandbox/**, Dockerfiles | Native ubuntu-24.04-arm runner: builds sandbox-runtime for linux/arm64 and smoke-renders one second at 320x180 through the stable headless-shell path. The docker job's sandbox-runtime entry also builds linux/amd64,linux/arm64 (arm64 under QEMU) |
infra-lint | the plumbing | infra/scripts/check-secrets.sh literal-secret tripwire (secrets.md), actionlint (+ shellcheck on run: blocks), ShellCheck on *.sh, hadolint on all eight Dockerfiles, docker compose config schema check |
tofu-lint | infra/** | OpenTofu 1.13.1: tofu fmt -check -recursive, tofu init -backend=false + tofu validate for envs/staging and envs/prod (covers every module they call; no credentials needed), tflint with the bundled terraform ruleset (infra/tofu/.tflint.hcl). Run locally: see infra/README.md. Required by the ci aggregate |
ci | — | Aggregate gate. Point branch protection at this single check rather than the matrix legs |
Path filtering
The changes job (dorny/paths-filter) decides which stacks a pull request
touched, so a copy tweak under /site does not rebuild eight container images.
Push events to master ignore the filter and run everything.
paths-filter lists the PR's files through the GitHub API, so the changes
job carries its own permissions: { contents: read, pull-requests: read } on
top of the workflow-wide contents: read. That is enough for Dependabot PRs
too: their runs get a read-only GITHUB_TOKEN scoped to the permissions the
workflow declares, and a read of the PR's file list is all the job needs.
Dependabot pull requests
Dependabot PRs run the same workflow as everyone else's and are judged by the
same ci check. Two things are worth knowing when one is red:
- Every job skipped,
Detect changed stacksfailed in under two seconds with no steps — the runner never started. Until October 2026 the repo's Dependabot PRs (#138–#154) all failed exactly like this, and the job's annotation said why:The job was not started because recent account payments have failed or your spending limit needs to be increased. That was a GitHub billing outage, not a permissions problem; nothing inci.ymlneeded to change. Dependabot PRs opened after billing was restored run and pass. Dependabot does not re-run CI on an existing PR by itself — comment@dependabot rebaseon it (or mergemasterinto it) to get a fresh run. - A
Docker · sandbox-*job fails onnpm installwithERESOLVE— the bump breaks a peer-dependency contract insandbox/template/package.json, which the runtime image installs without a lockfile. The React major is the usual culprit:@react-three/fiber9 and@react-three/drei10 need React 19, andreact-domcan never move withoutreact.dependabot.ymlgroups those packages so they are at least proposed together; the bump itself is a deliberate wave, since the template is what model-authored Remotion code renders against.
Dependabot's own security update jobs (the npm_and_yarn in /docs-site …
runs in the Actions list) are separate from CI. They fail with
security_update_not_possible when the vulnerable package is only reachable
through a dependency Dependabot is not allowed to bump; listing the directory
in dependabot.yml so the parent gets regular version updates is the fix.
Config files
| File | Purpose |
|---|---|
.golangci.yml | Shared by both Go modules — golangci-lint walks up from its working directory to find it. max-issues-per-linter/max-same-issues are 0 so CI never silently truncates findings |
.hadolint.yaml | failure-threshold: warning. DL3018/DL3008 (apk/apt version pinning) are off by design; reproducibility comes from the base image tag, not from distro package pins |
.github/dependabot.yml | Weekly updates for Actions, both Go modules, NuGet, the npm apps (web, site, docs-site, sandbox/template), and every Dockerfile base image. Packages that only work as a set are grouped so Dependabot proposes them as one PR: the Microsoft.Extensions.*/Microsoft.Agents.*/OpenAI/Azure.AI.OpenAI/Anthropic NuGet set, @mantine/*, react+react-dom+their types (in every npm directory), remotion+@remotion/*, @react-three/*, @docusaurus/* |
Running the same checks locally
# Go — all modules
for m in api sandbox runner; do
(cd $m && gofmt -l . && go vet ./... && go build ./... && go test ./... -race)
done
golangci-lint run ./... # from inside api/, sandbox/ or runner/
# .NET
dotnet build inference/ReelBolt.sln -c Release
dotnet test inference/ReelBolt.sln -c Release --no-build
# Next.js apps
for a in web site; do
(cd $a && npm ci && npx tsc --noEmit && npm run lint && npm run build)
done
# Infra
actionlint
shellcheck --severity=warning nginx/docker-entrypoint.sh caddy/docker-entrypoint.sh
hadolint --config .hadolint.yaml **/Dockerfile
docker compose config --quiet # needs JWT_SIGNING_KEY + SANDBOX_API_TOKEN set
Deployment
Slice S1 runs on two Hetzner VMs with docker compose, not Kubernetes (the topology is in
infra/README.md).
- Images. CI pushes every image tagged with the commit SHA (see the
dockerrow above). - Staging deploys automatically when CI succeeds on a
masterpush (workflow_run). Production isworkflow_dispatch(environmentproduction, optionalsha, default the tip ofmaster) and waits for that environment's required reviewers. - Order. The job verifies the images exist and their cosign signatures, then ssh-es to the
control VM, checks the repo out at the SHA, and runs
infra/compose/host.sh deploy <sha>(docker compose pull+up -d --waitwithinfra/compose/docker-compose.cloud.ymlplus the role file). Migrations are not a separate step: the Inference API migrates on startup and the engine after it, so deploying the control VM first and waiting for its health check is the API-then-engine order compose'sdepends_ongives locally. The engine VM follows. - Health and rollback.
/health,/api/v1/healthand/api/v1/workflow-engine/healthare probed through nginx. On a failed check each touched VM is moved back to the SHA recorded in/opt/reelbolt/.deploy.prev. A rollback restores code, not schema, so migrations must be backward compatible (expand, then contract). - Marketing site.
npm run build:pagesinsite/, thenwrangler pages deploy(staging to thestagingpreview branch, production tomain). Independent of the VMs.
Environment secrets (per GitHub environment staging and production): SSH_PRIVATE_KEY,
SSH_KNOWN_HOSTS (pinned host keys of both VMs), CONTROL_HOST, ENGINE_HOST, and for Pages
CLOUDFLARE_API_TOKEN + CLOUDFLARE_ACCOUNT_ID. Variables: CF_PAGES_PROJECT, optional
DEPLOY_USER (default root), NEXT_PUBLIC_SITE_URL, NEXT_PUBLIC_GA_MEASUREMENT_ID. The VMs
pull private GHCR images with the run's own GITHUB_TOKEN (packages: read), so no registry
token is stored. The per-host .env (database, R2, broker, JWT secrets) is the environment secret
CONTROL_ENV_FILE / ENGINE_ENV_FILE: the workflow writes it to /opt/reelbolt/.env over ssh with
mode 0600 before each host's deploy, and host.sh check-env validates it before anything restarts.
When a secret is unset the file already on the VM is kept. See secrets.md.
Known gaps
api/has no tests.go test ./...passes trivially there ([no test files]); the gate exists so the first test added is enforced from then on.sandbox/hasmain_test.goand it passes.- No integration tests. Nothing in CI starts Postgres, RabbitMQ, or the object store (Garage).
The
dotnetjob runs the WorkflowEngine unit tests (EFCore.InMemory + Moq) only. - SSH reachability for deploys (decided in F15). GitHub-hosted runners have no stable
address, so the F1 firewall now opens port 22 to the world (
deploy_ssh_cidrs, default0.0.0.0/0and::/0, ininfra/tofu/modules/network) and the host is hardened in cloud-init instead: key-onlysshd(PasswordAuthentication no,PermitRootLogin prohibit-password,MaxAuthTries 3,MaxStartups 10:30:60) plus afail2bansshd jail (3 failures in 10 minutes, one-hour ban). The deploy key and pinned host keys (SSH_KNOWN_HOSTS) are the only credentials. Why this option: a self-hosted runner means running and patching a third machine for a two-VM beta, and allow-listing GitHub's published ranges (api.github.com/meta) is hundreds of CIDRs that rotate, beyond what a Hetzner firewall rule holds. Keeping port 22 avoids touchingdeploy.yml, since moving to another port only reduces log noise and adds a second place that must agree. To close it again later, setdeploy_ssh_cidrs = [](ssh then followsadmin_cidrs) and move deploys to a self-hosted runner or a tunnel. cloud-init runs once (user_datais inignore_changes), so a host created before this change needs the same two files applied by hand. - Engine upstream. Resolved in F15: nginx reads
UPSTREAM_WORKFLOW_ENGINE(and the other upstreams) from the environment; the control override sets it tohttp://$ENGINE_PRIVATE_IP:8080, so/api/v1/workflow-engine/*and deploy.yml's engine health probe work across the two VMs.ENGINE_PRIVATE_IPis now also required in the control VM's.env. - The workflows themselves have never run (no accounts, no secrets); they are linted with
actionlint and the cloud compose files validated with
docker compose config.