Skip to main content

Staging

WP F14. A production-shaped staging environment: the same OpenTofu tree at smaller sizes, a separate Cloudflare R2 bucket and AWS KMS key, a Paddle sandbox instead of live billing, synthetic test organizations, a front door that keeps an unreleased environment private, and the scripts that run the load/cost test (X1) and the go-live checklist (X2) against it.

Staging exists because later work packages validate on it: F5 (R2 smoke test), F6 (hosted ASR and embeddings), F9 (a trace across the whole path), F12 (a burst that scales and drains), H2 (the pricing page against a live backend) and X1/X2. Decision R7 in plans/saas-launch/01-schedule.md pulls it forward to wave 2 for that reason.

Nothing in this document has been run against real accounts. No cloud, Paddle or DNS credential existed when it was written; every script below was tested against a local fake (infra/scripts/tests/), and the nginx front door was verified in a container. The sections that need a live provider say so explicitly.

1. What staging is​

Cloudflare DNS + proxy, R2 buckets reelbolt-staging-* (EU jurisdiction)
Hetzner reelbolt-staging-control (nginx, web, site, docs, go-api, inference, mcp, rabbitmq, qdrant)
reelbolt-staging-engine (workflow-engine, sandbox-executor, sandbox-runtime)
Scaleway managed PostgreSQL 16, fr-par, no HA, IP allow-list = the two VMs
AWS one KMS key alias/reelbolt-staging-dp-ring wrapping the Data Protection key ring
Paddle sandbox, not live

It is the same OpenTofu composition as prod (infra/tofu/envs/staging); only the inputs differ. The names come from local.name = "reelbolt-staging", so the buckets, the KMS alias, the SSH key and the Postgres instance are all distinct from prod's — applying staging cannot touch a prod resource.

stagingprod
control VMcx33 (shared vCPU)cx33
engine VMcx43cx53 (or ccx33 if the benchmark says so)
PostgresDB-PLAY2-PICODB-POP2-2C-8G
VM delete/rebuild protection, Hetzner backupsoffon
R2 bucket / KMS aliasreelbolt-staging-*reelbolt-prod-*
Paddlesandboxlive
front doorHTTP basic authopen (the app's own auth is the door)
autoscaler node cap2set by F12
storage retention30 days (D19)30 days

Create it exactly like prod (§ "Plan and apply staging" in infra/README.md): tofu init -backend-config=backend.hcl, tofu plan -out=plan.bin, tofu apply plan.bin, with the reelbolt-r2-state profile.

2. The front door​

An unreleased environment must not be indexable or pokeable. Two supported doors; pick one per environment.

HTTP basic auth (the default, no extra account needed). Set both variables in the control VM's .env (BASIC_AUTH_USER, BASIC_AUTH_PASSWORD); nginx renders an htpasswd line at container start and challenges the whole origin:

BASIC_AUTH_USER=staging
BASIC_AUTH_PASSWORD=<openssl rand -base64 24>
BASIC_AUTH_REALM=ReelBolt staging # optional

Both or neither: setting only one makes the nginx container refuse to start rather than serve an environment the operator believes is protected. /health and the Paddle webhook /api/v1/billing/webhooks/paddle are exempt — the first because deploy.yml probes it from the VM, the second because Paddle cannot answer a challenge (its signature HMAC is the credential there).

Cloudflare Access. Leave both variables blank and put an Access application in front of the hostname. go-live-check.sh --expect-front-door cloudflare accepts either a Cloudflare login redirect or a 401/403.

Calling a staging origin from a script​

With the door on, the browser's own path applies: the Authorization header carries the basic credential, and the session token travels in the reelbolt_token cookie that nginx/auth.js turns back into a bearer header for the API. Every script in infra/scripts/ therefore takes --basic-auth USER:PASSWORD (or STAGING_BASIC_AUTH in the environment) and sends the token as a cookie rather than as a bearer header. Sending a bearer header with basic auth on does not work: nginx consumes the header, and the upstream gets the cookie-derived value.

Verify the door locally (no cloud account)​

docker build -t reelbolt-nginx nginx/ # or: docker run nginx:alpine + apk add
docker run -d --name tls-test -p 127.0.0.1:18080:80 \
-e BASIC_AUTH_USER=staging -e BASIC_AUTH_PASSWORD=staging-secret \
-e UPSTREAM_GO_API=http://host.docker.internal:18081 reelbolt-nginx
curl -s -o /dev/null -w '%{http_code}\n' http://127.0.0.1:18080/ # 401
curl -s -o /dev/null -w '%{http_code}\n' -u staging:staging-secret http://127.0.0.1:18080/ # 200
curl -s -o /dev/null -w '%{http_code}\n' http://127.0.0.1:18080/health # reaches the upstream, no 401

3. Seeded test organizations​

export ADMIN_EMAIL=... ADMIN_PASSWORD=... # the platform admin on staging
infra/scripts/staging-seed-orgs.sh --base-url https://app.staging.example.com \
--orgs 10 --basic-auth staging:<password> --out staging-accounts.json

It creates N synthetic users through POST /api/v1/admin/users (each gets a Personal workspace, which is the synthetic organization), clears the mustChangePassword flag the admin API sets (the Go API's middleware refuses every other route until it is cleared), reads each account's Personal org id from GET /api/v1/orgs and writes staging-accounts.json (mode 0600, gitignored — it holds generated passwords and never the admin's).

Idempotent: a user that already exists is reused with the password from the file. A user whose password is unknown (the file was lost, or the password was reset) is reported and skipped; pass --recreate-missing to delete and recreate it. Addresses default to the reserved .invalid TLD, so a live SMTP configuration can never mail a stranger.

Keep the accounts file outside infra/: infra/scripts/check-secrets.sh reads a generated password there as a literal secret.

4. Load and cost test (X1's harness, run here)​

F14 step 4 is that the load and cost test runs on this environment; the harness itself is WP X1's and lives in infra/loadtest/. It is not duplicated here — one load driver, one cost reporter, tested by its own x1-selftest.sh. Read that README before running it: it states, per acceptance bullet, what could be performed in this tree and what could not.

infra/loadtest/x1-run.sh --base-url https://app.staging.example.com \
--admin-email [email protected] --admin-password '<password>' \
--fixture ./path-to-a-code-fixture --invoice invoices/october.csv

Two things this environment has to supply for it to mean anything:

  • Reachable providers and a real database. The driver seeds its own run-scoped organizations (x1-load-<timestamp>), uploads a code fixture, applies a template mix and submits 50 concurrent executions; on a staging stack with no provider rows configured the runs fail and the report says so rather than proving anything about the platform.
  • Non-placeholder rate cards. Every rate card shipped today is a B0 estimate, so x1-cost.mjs detects that state and reports the ±5% lane BLOCKED instead of comparing a placeholder against a real invoice. B13 is what unblocks it.

The synthetic organizations in §3 are a different thing: they are stable, named, idempotent accounts for the go-live check and for manual work (a workspace to click through, a file search to try, a trace to follow), not per-run load fixtures. The seeder's accounts file is what go-live-check.sh --accounts reads.

5. Go-live checklist on staging (F14's acceptance, X2)​

infra/scripts/go-live-check.sh --base-url https://app.staging.example.com \
--site-url https://staging.reelbolt.example --basic-auth staging:<password> \
--env-file /tmp/control.env --accounts staging-accounts.json \
--expect-front-door basic --expect-signup disabled --expect-billing sandbox

--env-file is the control VM's /opt/reelbolt/.env (the same text that is stored as the CONTROL_ENV_FILE environment secret). It is read with sed, never sourced, and only BOOLEAN/ENUM values are ever printed.

outcomemeaning
PASSthe environment answered as the checklist requires
FAILit did not; the run exits non-zero
WARNit did not, and that is another WP's open item (printed, never fatal)
MANUALit cannot be checked from outside; fatal only under --strict

What it checks automatically:

  • TLS reachability, /health, /api/v1/health, /api/v1/workflow-engine/health
  • the http:// → https:// redirect (FORCE_HTTPS=true)
  • the front door (basic auth challenge, or a Cloudflare Access refusal)
  • reelbolt_token is HttpOnly and Secure on a real login
  • POST /api/v1/auth/signup is closed or open, matching the launch expectation (the probe sends an empty body, which the handler rejects before reading it, so no account is created)
  • the legal pages serve no unfilled [TOKEN] placeholder (site/lib/legal-placeholders.ts)
  • TENANCY_MODE=Cloud, COOKIE_SECURE=true, SIGNUP_ENABLED, PADDLE_ENVIRONMENT, BILLING_ENABLED, DATA_PROTECTION_STORE=Database, STARTUP_CLEANUP_PURGE_QUEUES=false (D21), BASIC_AUTH_USER, POOL_MAX_NODES

What it will never check, and why: a live Paddle purchase and refund, a restored backup (F13), budget alerts armed at four providers, SPF/DKIM/DMARC at the registrar, the lawyer's sign-off, a status page, and a TLS Labs grade A. Each needs an account or a human this repository does not have. They are listed as MANUAL with their owner, and --strict is the mode to run immediately before announcing — it turns a leftover MANUAL item into a failure, so "every go-live checklist item runs green" cannot be claimed with a check still unrun.

Also still open, and flagged rather than faked: no HSTS header is served yet (a WARN; F15 owns the TLS Labs grade that requires it), and POOL_MAX_NODES exists only once F12's autoscaler reads it — until then the 2-node cap lives in infra/tofu/envs/staging and its output (worker_pool_max_nodes).

6. Local tests (what CI runs)​

bash infra/scripts/tests/run-tests.sh

Both scripts are driven against infra/scripts/tests/fake_staging.py, a fake origin answering the same routes with the shapes the Go API really returns (the admin user API, the login and change-password pair, /orgs, the health paths, the legal pages), with knobs (--front-door none|cloudflare, --signup enabled, --no-hsts, --legal-placeholders) that make one check fail on purpose. The suite asserts the classification of each observation (PASS/FAIL/WARN/MANUAL), the seeder's idempotency, the accounts-file mode and shape, that a stale password fails the cookie check rather than passing it, and the argument handling. It is wired into the infra-lint job.

It is not evidence that a real staging environment is green. That is exactly what §5 is for, and it needs the accounts.

7. Cost controls on staging​

  • The 2-node autoscaler cap (§1) and POOL_MAX_NODES.
  • The credit gate: a run whose balance cannot cover it is refused with a 402-class answer before any execution starts (B7/B6b), which infra/loadtest/x1-load.mjs --zero-credit-orgs exercises deliberately.
  • The kill switches (F11) are the first thing to reach for during a runaway test. They are not built yet, so today the lever is stopping the engine stack on the VM (host.sh stop in /opt/reelbolt) or pausing the render pool by hand.
  • The provider budget alerts in infra/README.md § Prerequisites are set on the account, not per environment: staging and prod share them.

8. Tear-down and rebuild​

Staging is disposable: tofu destroy in infra/tofu/envs/staging removes the VMs, the Postgres instance, the R2 buckets and the KMS key (the key has a 30-day deletion window). The state bucket and the GitHub environments' secrets are not managed here. After a rebuild, re-seed the test organizations (§3) — --recreate-missing is what makes that a one-liner when an old accounts file survives.