Secrets management
Operator reference for every secret the cloud deployment uses: what it is, how it reaches a VM,
and how to rotate it. The goal is that no secret lives in a manifest, in the repository or in
OpenTofu state. Slice S1 runs on two VMs with docker compose (decision D19), so delivery is
designed for VMs first; the Kubernetes plan for S2 is the last section. For the storage side see
cloud-storage.md; for the VM layout see
infra/README.md and
ci.md.
Rules
- A secret is a value whose disclosure lets someone spend money, read customer data or act as the platform. Identifiers (an AWS KMS key id, a bucket name, a hostname) are configuration, not secrets, but they travel in the same file for simplicity.
- Secrets are never committed, never passed as build arguments, never in an image, never in
OpenTofu variables with a default, and never in an output. The repository holds names and
placeholders only (
infra/compose/env.*.example). - The single source of truth for each secret's value is your password manager. GitHub environment secrets are write-only (you can replace a value, not read it back), so they are a delivery channel, not a vault.
- Secrets reach the VMs only over ssh, with file mode 0600, owned by the deploy user.
- Never log a value.
host.sh check-envreports variable names only. - The CI step
infra/scripts/check-secrets.shfails a build when infra, compose or workflow files contain what looks like a literal secret (see "Literal-secret check").
Secret inventory
Names are the .env variable names; the application setting each maps to is in
docker-compose.yml. "Both" means the control and engine hosts must hold the same value.
| Secret | Variable(s) | Used by | Host | Notes |
|---|---|---|---|---|
| JWT signing key | JWT_SIGNING_KEY | Go API (issues), Inference API and engine (validate; the engine also mints 5-minute on-behalf-of tokens), web | Both | HS256, 32+ characters. Single key, no overlap window: see "Rotating the JWT signing key". |
| Database credentials | DATABASE_CONNECTION_STRING (Npgsql), DATABASE_URL (Go API) | Inference API, engine, Go API | Control (both strings), engine (Npgsql string) | Two spellings of one database login on Scaleway managed PostgreSQL, sslmode=require. POSTGRES_PASSWORD is the self-host container's only. |
| Database admin password | TF_VAR_postgres_admin_password | OpenTofu only | Operator machine | Lands in encrypted state (the one unavoidable secret in state). Never reaches a VM. |
| RabbitMQ credentials | RABBITMQ_USER, RABBITMQ_PASSWORD | RabbitMQ container, Go API, Inference API, engine | Both | The broker listens on the control VM's private address. The broker applies the default user from these variables only when its volume is first created. |
| R2 object-store key pair | MINIO_ACCESS_KEY, MINIO_SECRET_KEY | Inference API, engine | Both | A bucket-scoped R2 token (cloud-storage.md "API tokens"). |
| Inference provider keys (environment) | ANTHROPIC_API_KEY, ANTHROPIC_AUTH_TOKEN, GEMINI_API_KEY, GOOGLE_API_KEY, DEEPSEEK_API_KEY, AZURE_OPENAI_API_KEY | Inference API, engine (only when a provider row's key is the literal env:, or for the legacy Azure fallback) | Both | Optional. Keys typed into the admin provider form are not here: they live encrypted in the database. A Pro/Max subscription token must not be used in a deployment that serves other people. |
| Inference provider keys (database) | stored per inference_providers row | Resolved by both services | Database | Encrypted with ASP.NET Data Protection. Rotated in the admin UI, see "Rotating inference provider keys". |
| AWS credentials for KMS | AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, AWS_REGION | Inference API and engine, to wrap and unwrap the Data Protection key ring | Both | Access key of the *-kms-wrapper IAM user created by the kms module (no key is created in OpenTofu). |
| Data Protection KMS key id | DATA_PROTECTION_KMS_KEY_ID | Same | Both | An identifier, not a secret. Required unless DATA_PROTECTION_KMS_LOCAL_KEY is set. |
| Data Protection local wrapper key | DATA_PROTECTION_KMS_LOCAL_KEY | Same | Self-host only | Base64, 32 bytes. Not for the cloud; see "Rotating the Data Protection wrapper". |
| Assistant confirmation-token key | none (stored in the platform_secrets table, wrapped as above) | Inference API | Database | One key per deployment, generated on first use. No operator action. |
| SMTP credentials | SMTP_USERNAME, SMTP_PASSWORD | Go API | Control | Optional. Without SMTP the API logs action links only in development. |
| Initial admin password | ADMIN_PASSWORD | Go API, first boot with an empty database | Control | Needed once. Delete it from the file afterwards. |
| Sandbox executor token | SANDBOX_API_TOKEN | Engine and sandbox executor (same host) | Engine | Required while S1 runs the sandbox on the engine VM. Self-host and S1 only; it goes away with the runner fabric. |
| Qdrant API key | VECTOR_QDRANT_API_KEY | Inference API | Control | Only when using hosted Qdrant (WP F7). The bundled container has no key and is not published. |
| OTLP export credentials | OTEL_EXPORTER_OTLP_HEADERS (and the non-secret OTEL_EXPORTER_OTLP_ENDPOINT, OTEL_EXPORTER_OTLP_PROTOCOL) | Go API, Inference API, engine | Both | Optional. The header carries the Grafana Cloud Authorization=Basic ... value (instance id plus a token scoped to metrics and traces write), so it is a secret; blank endpoint means nothing is exported. See observability.md. |
| Paddle billing credentials | PADDLE_API_KEY (config Billing__Paddle__ApiKey), PADDLE_WEBHOOK_SECRET (Billing__Paddle__WebhookSecret); non-secret BILLING_ENABLED, PADDLE_ENVIRONMENT, PADDLE_TOPUP_PRICE_500/_1000/_3000 | Inference API | Control | Only with BILLING_ENABLED=true; host.sh check-env then requires both secrets. The API key is a sandbox (pdl_sdbx_apikey_...) or live key and only works against its own environment. The webhook secret is the notification destination's secret key: Paddle signs every webhook with it, and anyone holding it can forge billing events, so rotate it in Paddle (it supports two active secrets during rotation) and then here. See billing.md. |
| Contact webhook URL | CONTACT_WEBHOOK_URL | Marketing site (site) | Control | A webhook URL usually embeds its token, so treat it as a secret. |
| Deploy ssh key | SSH_PRIVATE_KEY | deploy.yml | GitHub environment | Plus SSH_KNOWN_HOSTS (public host keys, pinned) and CONTROL_HOST/ENGINE_HOST. |
| rclone remote credentials | the file /opt/reelbolt/rclone.conf (mode 0600; no environment variable) | infra/scripts/backup-blobs.sh, via the reelbolt-backup.timer/reelbolt-backup.service units | Control | WP F13. The rclone config holding the credentials of the remotes it syncs to and from: the primary bucket's remote and the second (backup) provider's bucket — the nightly off-provider copy of the irreplaceable objects. It never reaches a container's environment or the .env files: the timer's BLOB_BACKUP_SOURCE/BLOB_BACKUP_TARGET variables carry rclone paths, not the remotes. Rotated out of band: replace the file on the control host (mode 0600, owned by the deploy user) with a fresh rclone config at the providers, then the next nightly run (or systemctl start reelbolt-backup.service) picks it up and the old credentials are deleted at the providers. |
| Environment files | CONTROL_ENV_FILE, ENGINE_ENV_FILE | deploy.yml writes them to /opt/reelbolt/.env | GitHub environment | The delivery channel for everything marked Control or Engine above. |
| Cloudflare Pages token | CLOUDFLARE_API_TOKEN, CLOUDFLARE_ACCOUNT_ID | deploy.yml (marketing site) | GitHub environment | Scope: Pages edit on the one project. |
| Infrastructure credentials | HCLOUD_TOKEN, SCW_ACCESS_KEY, SCW_SECRET_KEY, CLOUDFLARE_API_TOKEN (R2 and DNS edit), AWS_*, the reelbolt-r2 state profile, TF_ENCRYPTION passphrase | OpenTofu | Operator machine | Never reach a VM or GitHub. Scopes in infra/README.md. |
| Registry pull | GitHub Actions GITHUB_TOKEN | VMs, per deploy | Ephemeral | Minted for each run with packages: read; no registry token is stored. |
Declared by later work packages and not in the code yet. Each adds its name here and to
check_env in infra/compose/host.sh when it lands:
| Secret | Work package | Note |
|---|---|---|
| Runner pool credential | D (gateway), cloud workers | The pool worker's hello credential (docs/runner-protocol.md). |
| Gateway job-token key | D (gateway) | Signs per-job tokens. |
| Engine-to-gateway service token | D (gateway) | Protects the gateway's internal API. |
TURNSTILE_SECRET | C | Optional bot check on signup. |
| OAuth client secrets | C10 | Google and GitHub sign-in, later. |
How secrets reach a VM
Decision (S1): one GitHub environment secret per host holds that host's whole .env, and
deploy.yml writes it to /opt/reelbolt/.env over ssh with mode 0600 before each deploy.
The two secrets are CONTROL_ENV_FILE and ENGINE_ENV_FILE, defined per GitHub environment
(staging and production), next to the existing SSH_PRIVATE_KEY.
Why this and not a SOPS-encrypted file in the repository:
- GitHub is already the trust root for deploys. The ssh key that can reach both VMs lives in the same environment secrets, so storing the env files there adds no new party and no new key to lose or leak.
- Production's required reviewers already gate the
productionenvironment's secrets, so a pull request cannot read or change them. A SOPS file would need its own key (age or KMS) distributed to CI, and an encrypted blob in the repository is still a blob forever in history. - There is no new tooling: a rotation is "replace one secret, re-run the deploy". With SOPS it is a key ceremony plus a commit.
- The cost is real and accepted: GitHub secrets are write-only, so the password manager must hold the canonical file, and there is no review of a secret change in a pull request. Both are tolerable for one operator and two hosts. The S2 design (below) replaces this channel.
How it works (see .github/workflows/deploy.yml and infra/compose/host.sh):
- Before a host's deploy step, the workflow pipes the secret to
umask 077 && cat > .env.new, copies the previous file to.env.bak, and renames the new one into place. The write is atomic and the file is never world-readable, even for an instant. host.sh deployrunscheck_envfirst, before it pulls an image or touches a container. It requires mode 0600 (and fixes it), LF line endings, every required name present, noCHANGE_MEplaceholder, no development default for the RabbitMQ and object-store secrets, a JWT key of 32+ characters, and a Data Protection wrapper (KMS key id or local key) with AWS credentials when KMS is used. A failure exits non-zero and the running stack is untouched.docker compose up -dthen recreates exactly the containers whose environment changed.- If the secret is unset, the workflow keeps the file already on the VM, still validates it, and prints a notice. A host set up by hand keeps working.
First-time setup, per environment:
- Copy
infra/compose/env.control.exampleandenv.engine.exampleto your password manager (not into the repository) and fill everyCHANGE_ME. Generate secrets withopenssl rand -hex 32. The values marked "same as control" must match in both files. - Paste each filled file into the GitHub environment as
CONTROL_ENV_FILEandENGINE_ENV_FILE. - Run the deploy. To validate without deploying, ssh to a host and run
REELBOLT_ROLE=control bash /opt/reelbolt/infra/compose/host.sh check-env(roleengineon the engine host).
Adding or changing one variable: edit the canonical file in your password manager, paste the whole
file into the GitHub secret, re-run the deploy workflow for that environment. To roll back a bad
file, copy /opt/reelbolt/.env.bak over /opt/reelbolt/.env on the host and run
host.sh deploy <current sha>.
Escape hatch: scp a file as /opt/reelbolt/.env with mode 0600 and leave the GitHub secret
unset. This is the pre-F8 behaviour and what the first boot of a fresh VM uses.
Rotation basics
Every runbook below follows the same shape: create the new value, put it in the canonical file and the GitHub secret(s) for both environments' hosts that use it, run the deploy, verify, then revoke the old value. Where the dependency allows two live values at once the rotation has no downtime; where it does not, the section says so.
To roll a value, edit CONTROL_ENV_FILE (and ENGINE_ENV_FILE when the secret is shared),
then run Deploy for that environment from the Actions tab (a dispatch with an empty sha
redeploys the tip of master; a repeated SHA still recreates containers whose environment
changed). Check /health, /api/v1/health and /api/v1/workflow-engine/health afterwards;
the deploy workflow already does and rolls code back on failure.
No rotation drill has been run on a live staging environment yet (no cloud account exists at the time of writing). Run each section once on staging after the first deploy, and record the date in the status notes.
Rotating the JWT signing key
There is no overlap window today. The Go API signs with and validates against one key
(JWTSigningKey in api/services/jwt_service.go), and the Inference API, the engine
(IssuerSigningKey is a single SymmetricSecurityKey) and the web proxy do the same. A new key
therefore invalidates every issued token the moment a service restarts with it.
What that costs: every logged-in user is signed out (tokens last 24 hours) and must log in again; an execution that is mid-call when the engine and the API briefly disagree gets a 401 on its on-behalf-of token and retries; there is a window of a few minutes between the control and the engine deploys.
Runbook, at a quiet time:
openssl rand -hex 32; store it in the password manager.- Put it in
JWT_SIGNING_KEYin both env files and both GitHub secrets. - Run the deploy (control first, then engine, which
deploy.ymlalready orders). - Confirm a fresh login works and an execution started after the deploy completes.
- Tell users nothing; they simply sign in again.
Follow-up (not built): a zero-downtime rotation needs a second, validate-only key.
Concretely: Jwt__PreviousSigningKey / JWT_PREVIOUS_SIGNING_KEY, accepted by
TokenValidationParameters.IssuerSigningKeys in both .NET services and tried second by the Go
validator, never used to sign. The rotation then becomes: add the new key as primary and the old
as previous, deploy, wait 24 hours (the token lifetime), remove the previous key. Until that ships,
the table above is the real procedure.
Rotating the database password
Scaleway managed PostgreSQL allows several users, which gives a zero-downtime rotation.
- In the Scaleway console create a second database user with the same privileges on the
application database, with a new password from
openssl rand -base64 24. - Put the new user and password into
DATABASE_CONNECTION_STRINGandDATABASE_URL(control) andDATABASE_CONNECTION_STRING(engine); keepsslmode=require. - Run the deploy, then check the three health endpoints.
- Delete the old database user in the console after the deploy is healthy.
The admin password that OpenTofu manages (TF_VAR_postgres_admin_password) is separate and never
on a VM. To change it, edit your exported value and run tofu apply for the environment. The
state is encrypted, and the old password stays visible only in older state versions, so rotate it
if a state backup has ever been exposed.
Rotating the RabbitMQ password
The broker container creates its default user from RABBITMQ_USER/RABBITMQ_PASSWORD only on
the first start of an empty volume, so editing the variable alone does not change the broker.
There is a short interruption: services reconnect once they restart with the new value.
- Generate a value with
openssl rand -hex 24. - On the control VM:
docker compose --env-file .env --env-file .deploy.env -f docker-compose.yml -f infra/compose/docker-compose.cloud.yml -f infra/compose/docker-compose.cloud.control.yml exec rabbitmq rabbitmqctl change_password <user> '<new password>'. - Immediately put the value in
RABBITMQ_PASSWORDin both env files and run the deploy, so the consumers pick it up. MassTransit retries the connection until then. - Verify that an execution queues and completes.
Rotating the R2 object-store keys
No downtime. The application token is separate from every other credential (see
cloud-storage.md "API tokens").
- In Cloudflare create a second R2 token with the same scope (Object Read and Write, the one bucket).
- Put the new pair in
MINIO_ACCESS_KEYandMINIO_SECRET_KEYin both env files, deploy. - Run the staging smoke checklist in
cloud-storage.md, steps 1 to 5. - Wait at least 30 minutes: presigned URLs minted with the old key live up to 30 minutes and stop working the moment the old token is deleted.
- Delete the old token in Cloudflare.
Rotating inference provider keys
Keys stored in the admin UI (inference_providers rows). In /app/admin/inference-providers
open the row, paste the new key and save; the old one is overwritten (an omitted key means
"unchanged", an empty one clears it). The resolver caches rows for 60 seconds
(Inference:ProviderCacheSeconds), so the change applies within about a minute with no deploy.
Revoke the old key at the vendor afterwards.
Keys in the environment (ANTHROPIC_API_KEY, GEMINI_API_KEY, DEEPSEEK_API_KEY,
AZURE_OPENAI_API_KEY, used through the env: sentinel): create a new key at the vendor, put it
in both env files, deploy, run the provider's connection test in the admin UI, revoke the old key.
Rotating the AWS KMS access key and the Data Protection key ring
The IAM access key (AWS_ACCESS_KEY_ID/AWS_SECRET_ACCESS_KEY): an IAM user may hold two
access keys, so there is no downtime.
- In the AWS console create a second access key for the
*-kms-wrapperuser. - Put it in both env files, deploy.
- Confirm both services start (a wrong key fails at the first key-ring read) and a provider's connection test passes.
- Delete the old access key.
The KMS key itself rotates automatically each year (enable_key_rotation = true in the kms
module). AWS keeps the old key material, so existing wrapped keys keep decrypting and no
re-wrap is needed. Never delete the KMS key or schedule it for deletion: every stored provider
API key becomes unreadable (the same consequence as renaming the Data Protection application
name or the ReelForge.InferenceProvider.ApiKey purpose, decision D20).
Rotating the Data Protection wrapper
DATA_PROTECTION_KMS_LOCAL_KEY is for self-host only. It wraps the key ring with a 32-byte local
key. There is no re-wrap tool: replacing the value makes the existing wrapped keys undecryptable,
and with them every stored provider key. Do not rotate it. If it must change, first export the
provider keys from the admin UI (or re-enter them afterwards), and expect to re-enter all of them.
This is a known limit and a follow-up (a re-wrap command), and the cloud avoids it by using KMS.
Rotating the SMTP, Qdrant and webhook secrets
Each is an independent value held at a third party; none is shared across hosts.
- SMTP: create new credentials at the mail provider, set
SMTP_USERNAME/SMTP_PASSWORDin the control env file, deploy, send a test (an invite or a password reset), revoke the old credentials. - Qdrant API key (hosted Qdrant only): create a second key in the Qdrant Cloud console,
set
VECTOR_QDRANT_API_KEYin the control env file, deploy, check that file search returns results (notindexNotReady), delete the old key. - Contact webhook: regenerate the webhook URL at the receiving service, set
CONTACT_WEBHOOK_URL, deploy the control host, submit the contact form once.
Rotating the sandbox executor token
SANDBOX_API_TOKEN is shared by the engine and the sandbox executor, both on the engine VM, so a
single host changes. Active sandbox sessions end when the containers restart, so do it between
runs. Put a new openssl rand -hex 32 value in ENGINE_ENV_FILE, deploy, and run one execution
that uses Remotion. Self-host: change .env and docker compose up -d.
Rotating the deploy credentials
SSH_PRIVATE_KEY: generate a new key pair (ssh-keygen -t ed25519), add the new public key toauthorized_keyson both VMs, replace the GitHub secret, run a deploy, then remove the old public key from both VMs.SSH_KNOWN_HOSTS: only changes when a VM is rebuilt; regenerate withssh-keyscan -t ed25519against each host and verify the fingerprints out of band.CLOUDFLARE_API_TOKEN(Pages): roll the token in the Cloudflare dashboard and replace the GitHub secret; the next Pages deploy proves it.- Infrastructure credentials (
HCLOUD_TOKEN,SCW_*, the AWS key for OpenTofu, the R2 state-bucket token, the Cloudflare token fortofu): create the replacement at the provider, update your shell or~/.aws/credentials, runtofu planto prove it, delete the old one. These are never in GitHub or on a VM.
Rotating the admin bootstrap password
ADMIN_PASSWORD seeds the first admin only when the user table is empty, and a password change
is forced on first login. After the first boot, delete the line from the env file (and the GitHub
secret) and deploy. Rotate the admin's real password in the application.
Literal-secret check
infra/scripts/check-secrets.sh runs in the infra-lint job of ci.yml and scans infra/,
.github/workflows/ and docker-compose.yml. It fails on well-known token shapes (AWS access
key ids, GitHub tokens, sk-ant-, private key blocks and similar), on credentials embedded in a
URL, and on a secret-named assignment with a long opaque value. Placeholders pass: ${VAR},
<angle>, CHANGE_ME, var.x, secrets.X. Run it locally with
bash infra/scripts/check-secrets.sh. A reviewed false positive carries the comment
secret-scan: allow on that line. It is a tripwire for mistakes, not a replacement for review.
External Secrets Operator plan for S2
S2 moves the control plane to Kapsule (docs/design/cloud-provider-adr.md). The VM channel above
does not carry over; the target is External Secrets Operator (ESO) syncing from the provider's
secret manager into Kubernetes Secret objects. Nothing of this is built in S1.
- Store. Scaleway Secret Manager, because Kapsule and managed PostgreSQL are already Scaleway. ESO has a Scaleway provider; confirm the provider's current status at S2 time. Each variable in the inventory above becomes one secret entry named after the variable, per environment.
- Authentication. ESO authenticates to the secret manager with one scoped IAM API key
(read secrets in that project), created by hand and installed once as a bootstrap Kubernetes
Secret. It is the only secret applied outside ESO and is documented in the cluster runbook, not in a manifest. - Objects. A
ClusterSecretStoreand oneExternalSecretper workload that lists the same namescheck_envrequires today, withrefreshInterval: 1h, producing aSecretthe pods mount withenvFrom. The manifests live ininfra/k8s/external-secrets/and contain names and store references only, so the literal-secret check covers them unchanged. - Rotation. Update the entry in the secret manager; ESO syncs it; a reloader controller (or a checksum annotation on the deployment) restarts the pods. The per-secret runbooks above stay valid, with "edit the GitHub secret and deploy" replaced by "edit the secret-manager entry".
- Order of work. Move the secrets that have a zero-downtime path first (R2, provider keys, AWS key), then the shared ones (database, RabbitMQ), and the JWT key last, after dual-key validation exists.
Follow-ups
- Dual-key JWT validation (
Jwt__PreviousSigningKeyand the Go equivalent) so a rotation does not sign everyone out. Until then the JWT runbook above is the procedure. - A re-wrap command for
DATA_PROTECTION_KMS_LOCAL_KEYso self-host can rotate it. - Separate IAM users for the Inference API (Encrypt and Decrypt) and the engine (Decrypt only).
The
kmsmodule has one user and both hosts share its key today. - Run every rotation section once on staging and record the dates; none has run against real infrastructure yet.
- Add the not-yet-built secrets (gateway, Paddle, Turnstile, OAuth) to the inventory and to
check_envas their work packages land.