Skip to main content

Continuous Integration

CI builds and tests everything; a push to master also publishes signed images to GHCR, and a separate workflow deploys them to the cloud VMs and the marketing site to Cloudflare Pages (see "Deployment" below). Nothing deploys until the founder has created the GitHub environments and secrets: every deploy job skips with a notice when they are absent, so CI stays green before the accounts exist.

WorkflowFileTrigger
CI.github/workflows/ci.ymlpush to master, every PR, manual
Deploy.github/workflows/deploy.ymlgreen CI on master (staging), manual dispatch (production)
Runtime image release.github/workflows/runtime-image-release.ymlruntime-v* tag (the desktop runner's sandbox runtime image)
Runner release.github/workflows/runner-release.ymlrunner-v* tag (publishes the desktop runner to the beta channel), manual dispatch (promote to stable, or a build-and-verify dry run)

There is no CodeQL workflow. Code scanning on a private repository requires GitHub Code Security, which this repo does not have, so CodeQL could analyse but never upload and failed every run; it was removed rather than left red.

The runner release is the one workflow that does not follow the "skip with a notice until the accounts exist" rule the deploy jobs follow. It runs when somebody asks for a release, not on every push, so a release that silently does nothing is worse than one that names what is missing: its publish job fails with the checklist rather than skipping. The build legs still run, so the toolchain and the packaging are exercisable before the release key, the Windows certificate and the R2 credentials exist. See runner/build/README.md.

What runs​

JobCoversChecks
go (matrix: api, sandbox, runner)the three Go modulesgo mod tidy is a no-op, gofmt, go vet, go build ./..., go test ./... -race + coverage artifact
runner-windowsrunner/go vet and go test -race on windows-latest (runtime detection is OS-specific). The go job's runner leg also cross-compiles for windows/amd64 and linux/arm64
golangci-lint (matrix: api, sandbox, runner)the three Go modulesgolangci-lint v2.13.2, config in .golangci.yml
dotnetinference/ReelBolt.sln (Shared, Inference.Api, WorkflowEngine, WorkflowEngine.Tests)dotnet restore/build -c Release (Roslyn + .NET analyzers run here), dotnet test with TRX + Cobertura coverage artifacts
dotnet-formatsame solutiondotnet format --verify-no-changes --severity warn — blocking, and one of the jobs the ci aggregate requires
node (matrix: web, site)both Next.js appsnpm ci, tsc --noEmit, npm run lint (ESLint), npm test (Vitest, web only), npm run build
docs-sitedocs/ + docs-site/npm ci, npm run build (Docusaurus — fails on any broken link, so it doubles as the docs link check)
docker (matrix: 12 images, incl. mcp)every Dockerfilebuildx build with a GHA layer cache per image; on push to master only it also pushes ghcr.io/vecchiotom/reelbolt-<name>:<sha> and :master with an SBOM and provenance attestation and signs the digest with cosign (keyless, GitHub OIDC; packages: write/id-token: write are scoped to this job, PRs never push). Builds web/site/docs-site at their production target and sandbox at sandbox-runtime, control-plane, runner-gateway and worker (the D13 CloudPool worker, whose image is the only CI-built one that runs customer renders); inference-api and docs-site get the /docs tree through a named build-contexts entry
sandbox-runtime-arm64sandbox/**, DockerfilesNative ubuntu-24.04-arm runner: builds sandbox-runtime for linux/arm64 and smoke-renders one second at 320x180 through the stable headless-shell path. The docker job's sandbox-runtime entry also builds linux/amd64,linux/arm64 (arm64 under QEMU)
infra-lintthe plumbinginfra/scripts/check-secrets.sh literal-secret tripwire (secrets.md), actionlint (+ shellcheck on run: blocks), ShellCheck on *.sh, hadolint on all eight Dockerfiles, docker compose config schema check
tofu-lintinfra/**OpenTofu 1.13.1: tofu fmt -check -recursive, tofu init -backend=false + tofu validate for envs/staging and envs/prod (covers every module they call; no credentials needed), tflint with the bundled terraform ruleset (infra/tofu/.tflint.hcl). Run locally: see infra/README.md. Required by the ci aggregate
ci—Aggregate gate. Point branch protection at this single check rather than the matrix legs

Path filtering​

The changes job (dorny/paths-filter) decides which stacks a pull request touched, so a copy tweak under /site does not rebuild eight container images. Push events to master ignore the filter and run everything.

paths-filter lists the PR's files through the GitHub API, so the changes job carries its own permissions: { contents: read, pull-requests: read } on top of the workflow-wide contents: read. That is enough for Dependabot PRs too: their runs get a read-only GITHUB_TOKEN scoped to the permissions the workflow declares, and a read of the PR's file list is all the job needs.

Dependabot pull requests​

Dependabot PRs run the same workflow as everyone else's and are judged by the same ci check. Two things are worth knowing when one is red:

  • Every job skipped, Detect changed stacks failed in under two seconds with no steps — the runner never started. Until October 2026 the repo's Dependabot PRs (#138–#154) all failed exactly like this, and the job's annotation said why: The job was not started because recent account payments have failed or your spending limit needs to be increased. That was a GitHub billing outage, not a permissions problem; nothing in ci.yml needed to change. Dependabot PRs opened after billing was restored run and pass. Dependabot does not re-run CI on an existing PR by itself — comment @dependabot rebase on it (or merge master into it) to get a fresh run.
  • A Docker · sandbox-* job fails on npm install with ERESOLVE — the bump breaks a peer-dependency contract in sandbox/template/package.json, which the runtime image installs without a lockfile. The React major is the usual culprit: @react-three/fiber 9 and @react-three/drei 10 need React 19, and react-dom can never move without react. dependabot.yml groups those packages so they are at least proposed together; the bump itself is a deliberate wave, since the template is what model-authored Remotion code renders against.

Dependabot's own security update jobs (the npm_and_yarn in /docs-site … runs in the Actions list) are separate from CI. They fail with security_update_not_possible when the vulnerable package is only reachable through a dependency Dependabot is not allowed to bump; listing the directory in dependabot.yml so the parent gets regular version updates is the fix.

Config files​

FilePurpose
.golangci.ymlShared by both Go modules — golangci-lint walks up from its working directory to find it. max-issues-per-linter/max-same-issues are 0 so CI never silently truncates findings
.hadolint.yamlfailure-threshold: warning. DL3018/DL3008 (apk/apt version pinning) are off by design; reproducibility comes from the base image tag, not from distro package pins
.github/dependabot.ymlWeekly updates for Actions, both Go modules, NuGet, the npm apps (web, site, docs-site, sandbox/template), and every Dockerfile base image. Packages that only work as a set are grouped so Dependabot proposes them as one PR: the Microsoft.Extensions.*/Microsoft.Agents.*/OpenAI/Azure.AI.OpenAI/Anthropic NuGet set, @mantine/*, react+react-dom+their types (in every npm directory), remotion+@remotion/*, @react-three/*, @docusaurus/*

Running the same checks locally​

# Go — all modules
for m in api sandbox runner; do
(cd $m && gofmt -l . && go vet ./... && go build ./... && go test ./... -race)
done
golangci-lint run ./... # from inside api/, sandbox/ or runner/

# .NET
dotnet build inference/ReelBolt.sln -c Release
dotnet test inference/ReelBolt.sln -c Release --no-build

# Next.js apps
for a in web site; do
(cd $a && npm ci && npx tsc --noEmit && npm run lint && npm run build)
done

# Infra
actionlint
shellcheck --severity=warning nginx/docker-entrypoint.sh caddy/docker-entrypoint.sh
hadolint --config .hadolint.yaml **/Dockerfile
docker compose config --quiet # needs JWT_SIGNING_KEY + SANDBOX_API_TOKEN set

Deployment​

Slice S1 runs on two Hetzner VMs with docker compose, not Kubernetes (the topology is in infra/README.md).

  1. Images. CI pushes every image tagged with the commit SHA (see the docker row above).
  2. Staging deploys automatically when CI succeeds on a master push (workflow_run). Production is workflow_dispatch (environment production, optional sha, default the tip of master) and waits for that environment's required reviewers.
  3. Order. The job verifies the images exist and their cosign signatures, then ssh-es to the control VM, checks the repo out at the SHA, and runs infra/compose/host.sh deploy <sha> (docker compose pull + up -d --wait with infra/compose/docker-compose.cloud.yml plus the role file). Migrations are not a separate step: the Inference API migrates on startup and the engine after it, so deploying the control VM first and waiting for its health check is the API-then-engine order compose's depends_on gives locally. The engine VM follows.
  4. Health and rollback. /health, /api/v1/health and /api/v1/workflow-engine/health are probed through nginx. On a failed check each touched VM is moved back to the SHA recorded in /opt/reelbolt/.deploy.prev. A rollback restores code, not schema, so migrations must be backward compatible (expand, then contract).
  5. Marketing site. npm run build:pages in site/, then wrangler pages deploy (staging to the staging preview branch, production to main). Independent of the VMs.

Environment secrets (per GitHub environment staging and production): SSH_PRIVATE_KEY, SSH_KNOWN_HOSTS (pinned host keys of both VMs), CONTROL_HOST, ENGINE_HOST, and for Pages CLOUDFLARE_API_TOKEN + CLOUDFLARE_ACCOUNT_ID. Variables: CF_PAGES_PROJECT, optional DEPLOY_USER (default root), NEXT_PUBLIC_SITE_URL, NEXT_PUBLIC_GA_MEASUREMENT_ID. The VMs pull private GHCR images with the run's own GITHUB_TOKEN (packages: read), so no registry token is stored. The per-host .env (database, R2, broker, JWT secrets) is the environment secret CONTROL_ENV_FILE / ENGINE_ENV_FILE: the workflow writes it to /opt/reelbolt/.env over ssh with mode 0600 before each host's deploy, and host.sh check-env validates it before anything restarts. When a secret is unset the file already on the VM is kept. See secrets.md.

Known gaps​

  • api/ has no tests. go test ./... passes trivially there ([no test files]); the gate exists so the first test added is enforced from then on. sandbox/ has main_test.go and it passes.
  • No integration tests. Nothing in CI starts Postgres, RabbitMQ, or the object store (Garage). The dotnet job runs the WorkflowEngine unit tests (EFCore.InMemory + Moq) only.
  • SSH reachability for deploys (decided in F15). GitHub-hosted runners have no stable address, so the F1 firewall now opens port 22 to the world (deploy_ssh_cidrs, default 0.0.0.0/0 and ::/0, in infra/tofu/modules/network) and the host is hardened in cloud-init instead: key-only sshd (PasswordAuthentication no, PermitRootLogin prohibit-password, MaxAuthTries 3, MaxStartups 10:30:60) plus a fail2ban sshd jail (3 failures in 10 minutes, one-hour ban). The deploy key and pinned host keys (SSH_KNOWN_HOSTS) are the only credentials. Why this option: a self-hosted runner means running and patching a third machine for a two-VM beta, and allow-listing GitHub's published ranges (api.github.com/meta) is hundreds of CIDRs that rotate, beyond what a Hetzner firewall rule holds. Keeping port 22 avoids touching deploy.yml, since moving to another port only reduces log noise and adds a second place that must agree. To close it again later, set deploy_ssh_cidrs = [] (ssh then follows admin_cidrs) and move deploys to a self-hosted runner or a tunnel. cloud-init runs once (user_data is in ignore_changes), so a host created before this change needs the same two files applied by hand.
  • Engine upstream. Resolved in F15: nginx reads UPSTREAM_WORKFLOW_ENGINE (and the other upstreams) from the environment; the control override sets it to http://$ENGINE_PRIVATE_IP:8080, so /api/v1/workflow-engine/* and deploy.yml's engine health probe work across the two VMs. ENGINE_PRIVATE_IP is now also required in the control VM's .env.
  • The workflows themselves have never run (no accounts, no secrets); they are linted with actionlint and the cloud compose files validated with docker compose config.