Staging
WP F14. A production-shaped staging environment: the same OpenTofu tree at smaller sizes, a separate Cloudflare R2 bucket and AWS KMS key, a Paddle sandbox instead of live billing, synthetic test organizations, a front door that keeps an unreleased environment private, and the scripts that run the load/cost test (X1) and the go-live checklist (X2) against it.
Staging exists because later work packages validate on it: F5 (R2 smoke test), F6 (hosted
ASR and embeddings), F9 (a trace across the whole path), F12 (a burst that scales and
drains), H2 (the pricing page against a live backend) and X1/X2. Decision R7 in
plans/saas-launch/01-schedule.md pulls it forward
to wave 2 for that reason.
Nothing in this document has been run against real accounts. No cloud, Paddle or DNS
credential existed when it was written; every script below was tested against a local fake
(infra/scripts/tests/), and the nginx front door was verified in a container. The
sections that need a live provider say so explicitly.
1. What staging is
Cloudflare DNS + proxy, R2 buckets reelbolt-staging-* (EU jurisdiction)
Hetzner reelbolt-staging-control (nginx, web, site, docs, go-api, inference, mcp, rabbitmq, qdrant)
reelbolt-staging-engine (workflow-engine, sandbox-executor, sandbox-runtime)
Scaleway managed PostgreSQL 16, fr-par, no HA, IP allow-list = the two VMs
AWS one KMS key alias/reelbolt-staging-dp-ring wrapping the Data Protection key ring
Paddle sandbox, not live
It is the same OpenTofu composition as prod (infra/tofu/envs/staging); only the
inputs differ. The names come from local.name = "reelbolt-staging", so the buckets, the
KMS alias, the SSH key and the Postgres instance are all distinct from prod's — applying
staging cannot touch a prod resource.
| staging | prod | |
|---|---|---|
| control VM | cx33 (shared vCPU) | cx33 |
| engine VM | cx43 | cx53 (or ccx33 if the benchmark says so) |
| Postgres | DB-PLAY2-PICO | DB-POP2-2C-8G |
| VM delete/rebuild protection, Hetzner backups | off | on |
| R2 bucket / KMS alias | reelbolt-staging-* | reelbolt-prod-* |
| Paddle | sandbox | live |
| front door | HTTP basic auth | open (the app's own auth is the door) |
| autoscaler node cap | 2 | set by F12 |
| storage retention | 30 days (D19) | 30 days |
Create it exactly like prod (§ "Plan and apply staging" in
infra/README.md): tofu init -backend-config=backend.hcl,
tofu plan -out=plan.bin, tofu apply plan.bin, with the reelbolt-r2-state profile.
2. The front door
An unreleased environment must not be indexable or pokeable. Two supported doors; pick one per environment.
HTTP basic auth (the default, no extra account needed). Set both variables in the
control VM's .env (BASIC_AUTH_USER, BASIC_AUTH_PASSWORD); nginx renders an
htpasswd line at container start and challenges the whole origin:
BASIC_AUTH_USER=staging
BASIC_AUTH_PASSWORD=<openssl rand -base64 24>
BASIC_AUTH_REALM=ReelBolt staging # optional
Both or neither: setting only one makes the nginx container refuse to start rather than
serve an environment the operator believes is protected. /health and the Paddle webhook
/api/v1/billing/webhooks/paddle are exempt — the first because deploy.yml probes it
from the VM, the second because Paddle cannot answer a challenge (its signature HMAC is the
credential there).
Cloudflare Access. Leave both variables blank and put an Access application in front of
the hostname. go-live-check.sh --expect-front-door cloudflare accepts either a Cloudflare
login redirect or a 401/403.
Calling a staging origin from a script
With the door on, the browser's own path applies: the Authorization header carries the
basic credential, and the session token travels in the reelbolt_token cookie that
nginx/auth.js turns back into a bearer header for the API. Every script in
infra/scripts/ therefore takes --basic-auth USER:PASSWORD (or STAGING_BASIC_AUTH in
the environment) and sends the token as a cookie rather than as a bearer header. Sending a
bearer header with basic auth on does not work: nginx consumes the header, and the upstream
gets the cookie-derived value.
Verify the door locally (no cloud account)
docker build -t reelbolt-nginx nginx/ # or: docker run nginx:alpine + apk add
docker run -d --name tls-test -p 127.0.0.1:18080:80 \
-e BASIC_AUTH_USER=staging -e BASIC_AUTH_PASSWORD=staging-secret \
-e UPSTREAM_GO_API=http://host.docker.internal:18081 reelbolt-nginx
curl -s -o /dev/null -w '%{http_code}\n' http://127.0.0.1:18080/ # 401
curl -s -o /dev/null -w '%{http_code}\n' -u staging:staging-secret http://127.0.0.1:18080/ # 200
curl -s -o /dev/null -w '%{http_code}\n' http://127.0.0.1:18080/health # reaches the upstream, no 401
3. Seeded test organizations
export ADMIN_EMAIL=... ADMIN_PASSWORD=... # the platform admin on staging
infra/scripts/staging-seed-orgs.sh --base-url https://app.staging.example.com \
--orgs 10 --basic-auth staging:<password> --out staging-accounts.json
It creates N synthetic users through POST /api/v1/admin/users (each gets a Personal
workspace, which is the synthetic organization), clears the mustChangePassword flag the
admin API sets (the Go API's middleware refuses every other route until it is cleared),
reads each account's Personal org id from GET /api/v1/orgs and writes
staging-accounts.json (mode 0600, gitignored — it holds generated passwords and never
the admin's).
Idempotent: a user that already exists is reused with the password from the file. A
user whose password is unknown (the file was lost, or the password was reset) is reported
and skipped; pass --recreate-missing to delete and recreate it. Addresses default to the
reserved .invalid TLD, so a live SMTP configuration can never mail a stranger.
Keep the accounts file outside infra/: infra/scripts/check-secrets.sh reads a
generated password there as a literal secret.
4. Load and cost test (X1's harness, run here)
F14 step 4 is that the load and cost test runs on this environment; the harness itself
is WP X1's and lives in infra/loadtest/. It is not
duplicated here — one load driver, one cost reporter, tested by its own
x1-selftest.sh. Read that README before running it: it states, per acceptance bullet,
what could be performed in this tree and what could not.
infra/loadtest/x1-run.sh --base-url https://app.staging.example.com \
--fixture ./path-to-a-code-fixture --invoice invoices/october.csv
Two things this environment has to supply for it to mean anything:
- Reachable providers and a real database. The driver seeds its own run-scoped
organizations (
x1-load-<timestamp>), uploads a code fixture, applies a template mix and submits 50 concurrent executions; on a staging stack with no provider rows configured the runs fail and the report says so rather than proving anything about the platform. - Non-placeholder rate cards. Every rate card shipped today is a B0 estimate, so
x1-cost.mjsdetects that state and reports the ±5% lane BLOCKED instead of comparing a placeholder against a real invoice. B13 is what unblocks it.
The synthetic organizations in §3 are a different thing: they are stable, named, idempotent
accounts for the go-live check and for manual work (a workspace to click through, a file
search to try, a trace to follow), not per-run load fixtures. The seeder's accounts file is
what go-live-check.sh --accounts reads.
5. Go-live checklist on staging (F14's acceptance, X2)
infra/scripts/go-live-check.sh --base-url https://app.staging.example.com \
--site-url https://staging.reelbolt.example --basic-auth staging:<password> \
--env-file /tmp/control.env --accounts staging-accounts.json \
--expect-front-door basic --expect-signup disabled --expect-billing sandbox
--env-file is the control VM's /opt/reelbolt/.env (the same text that is stored as the
CONTROL_ENV_FILE environment secret). It is read with sed, never sourced, and only
BOOLEAN/ENUM values are ever printed.
| outcome | meaning |
|---|---|
PASS | the environment answered as the checklist requires |
FAIL | it did not; the run exits non-zero |
WARN | it did not, and that is another WP's open item (printed, never fatal) |
MANUAL | it cannot be checked from outside; fatal only under --strict |
What it checks automatically:
- TLS reachability,
/health,/api/v1/health,/api/v1/workflow-engine/health - the
http://→https://redirect (FORCE_HTTPS=true) - the front door (basic auth challenge, or a Cloudflare Access refusal)
reelbolt_tokenisHttpOnlyandSecureon a real loginPOST /api/v1/auth/signupis closed or open, matching the launch expectation (the probe sends an empty body, which the handler rejects before reading it, so no account is created)- the legal pages serve no unfilled
[TOKEN]placeholder (site/lib/legal-placeholders.ts) TENANCY_MODE=Cloud,COOKIE_SECURE=true,SIGNUP_ENABLED,PADDLE_ENVIRONMENT,BILLING_ENABLED,DATA_PROTECTION_STORE=Database,STARTUP_CLEANUP_PURGE_QUEUES=false(D21),BASIC_AUTH_USER,POOL_MAX_NODES
What it will never check, and why: a live Paddle purchase and refund, a restored
backup (F13), budget alerts armed at four providers, SPF/DKIM/DMARC at the registrar, the
lawyer's sign-off, a status page, and a TLS Labs grade A. Each needs an account or a human
this repository does not have. They are listed as MANUAL with their owner, and
--strict is the mode to run immediately before announcing — it turns a leftover
MANUAL item into a failure, so "every go-live checklist item runs green" cannot be
claimed with a check still unrun.
Also still open, and flagged rather than faked: no HSTS header is served yet (a WARN;
F15 owns the TLS Labs grade that requires it), and POOL_MAX_NODES exists only once F12's
autoscaler reads it — until then the 2-node cap lives in infra/tofu/envs/staging and its
output (worker_pool_max_nodes).
6. Local tests (what CI runs)
bash infra/scripts/tests/run-tests.sh
Both scripts are driven against infra/scripts/tests/fake_staging.py, a fake origin
answering the same routes with the shapes the Go API really returns (the admin user API, the
login and change-password pair, /orgs, the health paths, the legal pages), with knobs
(--front-door none|cloudflare, --signup enabled, --no-hsts, --legal-placeholders)
that make one check fail on purpose. The suite asserts the classification of each
observation (PASS/FAIL/WARN/MANUAL), the seeder's idempotency, the accounts-file mode and
shape, that a stale password fails the cookie check rather than passing it, and the argument
handling. It is wired into the infra-lint job.
It is not evidence that a real staging environment is green. That is exactly what §5 is for, and it needs the accounts.
7. Cost controls on staging
- The 2-node autoscaler cap (§1) and
POOL_MAX_NODES. - The credit gate: a run whose balance cannot cover it is refused with a 402-class answer
before any execution starts (B7/B6b), which
infra/loadtest/x1-load.mjs --zero-credit-orgsexercises deliberately. - The kill switches (F11) are the first thing to reach for during a runaway test. They are
not built yet, so today the lever is stopping the engine stack on the VM
(
host.sh stopin/opt/reelbolt) or pausing the render pool by hand. - The provider budget alerts in
infra/README.md§ Prerequisites are set on the account, not per environment: staging and prod share them.
8. Tear-down and rebuild
Staging is disposable: tofu destroy in infra/tofu/envs/staging removes the VMs, the
Postgres instance, the R2 buckets and the KMS key (the key has a 30-day deletion window).
The state bucket and the GitHub environments' secrets are not managed here. After a rebuild,
re-seed the test organizations (§3) — --recreate-missing is what makes that a one-liner
when an old accounts file survives.