Metering and billing (usage, credits, plans and the provider boundary)
ReelBolt meters every unit of work it buys from a provider and charges the customer for the
platform-supplied ones in credits. This document is the developer reference for that machinery:
the eleven usage kinds and how one billable unit becomes a usage_events row, the two numbers every
row carries (our USD cost and the credits charged), rate cards and rate multipliers (including what
happens to a platform unit nobody has priced yet), the credit ledger and its reserve → settle →
refund lifecycle with the failure classes that refund, the overdraft tolerance that lets a step finish
just over its estimate, entitlements and the plan catalog that is mirrored into the marketing site and
drift-checked in CI, the IBillingProvider boundary and the webhook pipeline with its idempotency
keys, and the subscription state machine.
The design record behind it is docs/design/saas-metering.md (usage taxonomy and failure
classification, WP B1) and docs/design/saas-pricing-and-billing.md (the credit unit, the top-up
question and the Paddle ADR, WP B0), with the product decisions D8–D12 in
plans/saas-launch/00-decisions.md. The code is inference/src/ReelBolt.Shared/Metering/ (recorder,
rate cards, the metering client decorators), inference/src/ReelBolt.Shared/Billing/ (ledger, plans,
entitlement mapping), inference/src/ReelBolt.WorkflowEngine/Services/Billing/CreditGate.cs (the
per-step gate) and inference/src/ReelBolt.Inference.Api/Services/Billing/ (webhooks, the provider
boundary, provisioning).
Provider configuration is not in this page. The Paddle environment variables, the plan price ids, the checkout/portal/seats endpoints, the operator setup checklist and the Paddle wire formats are all in docs/billing.md. This page links to it rather than repeating it.
Billing is off by default. Billing:Enabled=false (the self-host default) makes every entitlement
unlimited, leaves the credit gate a no-op, and 404s the billing endpoints — so an install with no
billing configured records usage but charges nobody.
Usage kinds: the eleven billable units
ReelBolt's usage taxonomy is the UsageKind enum in
inference/src/ReelBolt.Shared/Data/Models/Enums.cs: eleven append-only members, one per billable
physical unit. The enum is persisted by member name (EF string conversion) in usage_events.kind, so
appending a member needs no migration and renaming one breaks stored rows. usage_events.quantity is
a decimal(28,6) in the unit named below.
| Member | Unit | Produced by in ReelBolt |
|---|---|---|
LlmInputTokens | tokens | every chat call: workflow agents, ReviewLoop, room steps, the assistant, the file summarizer, decision and guardrail calls on logprob clients, vision captioning. The quantity is input net of the cached portion, so the three LLM kinds never overlap. |
LlmOutputTokens | tokens | the same calls; the model's output tokens. |
LlmCachedInputTokens | tokens | the part of the input the provider served from its prompt cache, billed at its own (lower) rate. Not to be confused with the engine's step-result cache (FromCache), which costs nothing and is a separate concept. |
EmbeddingTokens | tokens | file indexing, semantic file search queries, platform-docs indexing and per-turn doc retrieval. The provider's reported count, or ceil(chars / 4) when a backend reports none. |
AsrAudioSeconds | seconds of audio sent | VideoAnalyze transcription and voiceover word alignment. |
TtsInputBytes | UTF-8 bytes of text sent | the Voiceover step, through ISpeechSynthesisClient (Fish Audio and the OpenAI-compatible speech endpoint both bill per byte). |
VideoGenSeconds | seconds of generated video | the VideoGenerate step, from the take's requested seconds (the platform-side job spec), not the vendor's reported usage. |
RenderOutputSeconds | seconds of rendered Remotion output | the sandbox Remotion render. The engine fetches the stored object and probes it; the sandbox is not asked. |
EncodeOutputSeconds | seconds of ffmpeg-encoded output | VideoCompile, from the probe of the file it just encoded. |
AnalyzeMediaSeconds | seconds of source media analysed | the VideoAnalyze extraction phase: probe, silence and shot detection, frame grid, audio, keyframes. One row per analysed source. |
StorageGbDay | gigabyte-days stored | the daily storage sampler, one row per organization. |
Two columns beside the kind carry attribution that is deliberately not part of it:
usage_events.source(UsageSource) isWorkflow,Assistant,Indexing,DocsIndex,StorageorUnattributed, and says which subsystem produced the call.usage_events.billing_mode(UsageBillingMode) isPlatform(a platform-managed provider),Byo(the organization's own provider) orRunner(compute on the organization's own runner). The sameLlmInputTokensevent isPlatformorByodepending on who owns the resolved provider row —MeteringChatClient.BillingModeOfreads exactly that: a nullOrganizationIdon the provider meansPlatform, anything else meansByo.
BYO and runner work is still recorded — with the quantity and with credits_millis = 0 and
cost_usd = 0 — so a usage dashboard stays complete and an organization can see what its own key did,
even though none of it is charged.
How one billable unit becomes a usage_events row
Every metered call in ReelBolt goes through one interface,
IUsageRecorder.RecordAsync(UsageRecordRequest) (inference/src/ReelBolt.Shared/Metering/), and one
implementation, UsageRecorder. UsageRecorder is registered by AddReelBoltUsageRecorder<TContext>()
as a singleton and a hosted service: callers get the charged amount back synchronously while the
row itself is buffered and written in batches (100 rows or two seconds, whichever comes first; four
write attempts with a one-second delay between them).
A row is built in four steps:
- Attribution is ambient, not passed in.
UsageRecordRequestcarries only the kind, quantity, provider id and kind, model name and billing mode. Organization, project, execution, step result, user and source come fromUsageScope.Current, anAsyncLocalopened by the caller that owns the unit of work:WorkflowExecutorServiceopens one per step iteration (after theworkflow_step_resultsrow exists, so the scope carries itsStepResultId), andAssistantChatControlleropens one per assistant turn (Assistant), while file indexing and the platform-docs index open their own. A request may override the scope explicitly. A call that runs outside any scope is not lost: it is written against thePlatformorganization withsource = Unattributedand logged as a warning, so an unmetered code path shows up in the ledger and in the logs instead of silently costing money. - The quantity is what the provider or the job spec says, never an estimate of elapsed work. A provider that reports no usage gets the platform's own count (the UTF-8 byte length of the request text, or a character-based token estimate) rather than a silent zero.
- Pricing happens at write time through
IRateCardService.PriceAsync, and the resultingPriceQuotegives the row itscredits_millis,cost_usdandrate_card_id. Pricing a row when it is created — rather than when it is read — is what makes a later rate-card change apply to new usage only. - The row is queued on an unbounded channel, and the amount is added to the scope's running credits, which is what a mid-step check can read without touching the database.
Idempotency is the row's identity. Every row carries a unique idempotency_key (a unique index on
the column), so a retried write is a no-op. The default key is
{stepResultId or scopeId}:{callSeq}:{kindCode}, where callSeq is a per-scope counter and
kindCode is in, out or cin for the three LLM kinds and the enum name otherwise — one model
call therefore writes one row per kind under one shared call number. A unit with a natural identity of
its own overrides the key: a video-generation take is videogen:{jobId:N}, so a replayed settle never
double-charges, and the daily storage sample is storage:{org}:{yyyy-MM-dd}. The sink
(EfUsageEventSink<TContext>) checks the keys before inserting, and on a unique-index race falls back
to inserting row by row so one duplicate never loses its neighbours. If all four write attempts fail,
the batch's keys and amounts are written to the log at Critical level, explicitly so they can be
replayed by hand.
One call writes one row per kind, and the LLM split is exact. MeteringChatClient wraps every
client the chat factory hands out and, on each response, records LlmInputTokens = input - cached,
LlmOutputTokens = output and LlmCachedInputTokens = min(cached, input), using the same call number
for all three. Embedding, ASR, TTS and decision clients are wrapped the same way by their own
factories, which is the point of the design: a call site never has to remember to meter, because the
factory is the choke point. A streaming chat call reports its usage as the sum of the streamed
updates, and the record runs in a finally block, so a consumer that stops early still pays for the
tokens it caused.
Cost and credits: two numbers on every row
Every usage_events row carries both numbers (decision D10):
cost_usd— our provider cost (COGS) in USD for that quantity. It is the one that decides whether the business is making money, and it is the number the pricing work package B13 is meant to replace the placeholder estimates with. It is0forByoandRunnerrows, because that work is not our cost.credits_millis— the credits charged to the organization, in integer millicredits (1 credit = 1,000 millicredits, decision D8). Integers rather than a float because embedding and ASR charges are fractions of a credit (a 500-token embedding is well under one millicredit) and must accumulate without rounding to zero, and because abigintcannot drift.
The link between the two is the rate card: a card holds cost_usd_per_unit and
credits_millis_per_unit, and the seed derives the credit column from the USD column at
1 credit = $0.006 of COGS (the margin target agreed in the pricing record), so the credit price of a
unit is a data change rather than a code change.
Two rounding rules matter, and both are in RateCardService.Price: the charge is
round-half-up of quantity × credits_millis_per_unit × multiplier — rounded once per event, never
per unit — and a non-zero platform-billed event is floored at one millicredit, so metering never
reads as free. cost_usd is not rounded at all: it is stored at numeric(18,8).
Rate cards and rate multipliers
rate_cards and rate_multipliers (owned by the Inference API, mapped read-only by the
WorkflowEngine) are the data behind pricing; the Inference API exposes them on an admin surface, and
the engine reads them through the same RateCardService singleton, cached for 60 seconds.
A rate card prices one usage kind for one provider kind and one model pattern:
| Column | Meaning |
|---|---|
provider_kind + usage_kind + model_pattern + effective_from | The unique key of the row. |
model_pattern | An exact model name, or prefix*, or * for any model of that provider kind. The most specific match wins: an exact name beats the longest prefix, which beats *. |
credits_millis_per_unit / cost_usd_per_unit | The two prices of one unit. |
effective_from / effective_to | The card applies over [from, to). A card is superseded by dating, never by editing in place, so historical rows keep the price that was actually charged; between two cards of equal specificity the later effective_from wins. |
A rate multiplier scales the credits of a priced event by billing mode: one row per
(billing_mode, usage_kind) pair, where a null usage_kind is the default for the mode and a
kind-specific row beats it. No row at all means Platform is charged in full (1.0) and Byo and
Runner are free (0) — which is how decision D8's "BYO inference costs 0 credits" and "a desktop
runner costs 0 compute credits" are expressed as data rather than code. The seed writes exactly those
three rows: Platform = 1, Byo = 0, Runner = 0.
Compute is priced by the job spec, and its resolution tier rides in the model field. The three
compute kinds are priced from what the platform itself measured — the probed duration of the file it
rendered or encoded, or of the source it analysed — and never from wall-clock time or a
runner-reported figure, because a runner is untrusted (decisions D13–D16). The resolution tier is the
model_name on the row, so the existing model_pattern matching prices a tier with no schema change:
ComputeMetering.TierFor(width, height) takes the short side of the probed frame, so <= 720 is
720p, <= 1080 is 1080p and anything larger is uhd (a vertical 1080x1920 reel is 1080p, and
an unknown size such as audio-only prices as 720p). Compute rows are filed under the
OpenAICompatible provider kind, because rate_cards keys on a provider kind and compute has none; a
new InferenceProviderKind member would instead surface in every provider picker and allowlist.
The seed data is indicative, and it says so. RateCardSeed writes cards only while rate_cards
is empty. Every row carries a note that still begins with the marker word PLACEHOLDER — the admin
table badges on it and infra/loadtest/x1-cost.mjs keeps its invoice-reconciliation lane BLOCKED
while any active card carries it — followed by that row's own basis, because none of these figures has
been reconciled against a provider invoice. The credits themselves are not estimates: D9 is locked at
1 credit = $0.006 of COGS (a 70% gross margin at Pro's effective credit rate), so
credits_millis_per_unit = cost_usd_per_unit / 0.006 * 1000 for every row. What B13 replaces is the
USD cost side, which is where the written record runs out:
| Cards | Cost per unit | Basis |
|---|---|---|
| chat / vision / decision tokens, 7 provider kinds | $1.20 / $4.80 / $0.12 per million for input, output, cached input | B0's Mid scenario, $3.00 per million blended, split 1:4 over a 50/50 token mix so the blend lands exactly on $3.00/Mtok. B0 labels all three scenarios "estimated". |
| embeddings | $0.01 per million tokens | researched list-price class (B0's own worked example). |
| ASR | $0.006 per audio minute | researched list price, OpenAI whisper-1 / gpt-4o-transcribe (docs/design/cloud-provider-adr.md). |
| TTS | $15 per million bytes | genuinely unmeasured. B0 lists voiceover under "not measured", and no published per-byte price is recorded anywhere in this repo; the original figure is kept rather than replaced by an invented one. |
| video generation | $0.08 per second | MiniMax-H3-Max at its 768P list price (MiniMaxVideoGenerationClient.Describe), reconciling with the measured ledger ($9.20 over 18 succeeded jobs = $0.511 per clip at ~6.4 s). |
| compute (9 cards) | $0.000017 per 720p media second, x2.25 for 1080p and x9 for UHD, with encode at 0.5 and analyse at 0.25 of render | B0's researched anchor: an 8 vCPU VM at $0.25/hour with 2 vCPU effective per run is about $1.7e-5 per run-second. The run-second-to-output-second conversion is estimated (read at one-to-one, the conservative end). |
Two known gaps in the same area, stated rather than implied: a rate card keys on the model name and not the resolution, so the single wildcard video-generation card cannot price 480P ($0.05/s) and 768P ($0.08/s) apart; and chat tokens carry one mid-tier figure per provider kind rather than a per-model price, which is exactly what B13's measured data is for.
An unpriced platform unit is recorded at zero credits and logged, not thrown. Pricing itself fails
closed: RateCardService.Price throws UnpricedUsageException when billing_mode = Platform and no
card matches. UsageRecorder is where that becomes survivable: it catches the exception, logs it at
Error level ("no rate card; recorded unpriced"), and still writes the row — with
credits_millis = 0 and no rate_card_id — so the gap is visible in the ledger and loud enough to
page on, while a model call the customer has already paid for is never turned into a workflow failure.
For Byo and Runner a missing card is not an error at all: those return zero credits and zero cost
directly.
Two known gaps in the same area, stated rather than implied: the daily storage sampler writes its
StorageGbDay row directly, with credits_millis = 0 and cost_usd = 0 (its code carries a
TODO(B4) to price through the rate-card service), so storage is metered for visibility but charged
nothing; and compute is recorded on success only, so a failed render or encode is not metered at all
until the failure classes are wired (see "What is designed but not built yet" below).
The credit ledger: grants, holds and refunds
ReelBolt.Shared/Billing/CreditLedger.cs is the money. It is one class over four tables, and every
operation follows the same shape: one transaction, a per-organization
pg_advisory_xact_lock(hashtext(org_id)), an immutable double-entry credit_ledger_entries row, and
the credit_accounts projection updated in that same transaction — so the balance can never disagree
with the ledger. Every call takes an idempotency key; replaying a key returns the original entry and
changes nothing (and refuses if the key belongs to a different organization).
| Table | What it holds |
|---|---|
credit_accounts | One row per organization: balance_millis, reserved_millis (held by open reservations), and a row_version optimistic-concurrency token. Spendable now is balance - reserved. |
credit_grants | Buckets of credits with their own expiry and source (Trial, Plan, TopUp, Promo). |
credit_ledger_entries | The append-only journal. Kinds: Grant, TopUp, PlanRefill, Reserve, Settle, Release, Refund, Adjust, Expire. Each entry carries a balance delta and a reserved delta, the balance after it, an optional ref_entry_id (a settle or release points at its reserve, a refund at its settle) and, on a settle, an allocations_json array naming the grants it drew from. |
trial_claims | One free trial per normalized email (SHA-256 only, so deleting an account does not let the address claim again). |
The invariant the ledger maintains is sum(grant.remaining) == max(balance, 0), and
IsConsistentAsync is the audit check for it (plus "the balance equals the sum of the entries'
deltas"). Credits are consumed from the buckets in a fixed order: non-top-up before top-up, then
earliest expiry first, never-expiring last, then oldest — so plan credits burn before bought ones,
and a top-up is the last thing to go.
The five operations:
GrantAsyncadds a bucket, and absorbs a negative balance first (credits arriving over a debt repay it). The trial grant, the monthly plan refill and a top-up all use it.ReserveAsyncholds an amount before work starts, refusing withINSUFFICIENT_CREDITSunlessavailable - amount >= minAvailableAfterMillis. Expired grants are swept first, so they can never back a hold.SettleAsynccloses a reservation: it drops the hold and charges the actual amount, which may exceed the hold. The overage is charged and the balance may go negative; the floor policy lives in the caller (the gate), not the ledger.ReleaseAsyncdrops a hold without charging — a step that consumed nothing.RefundAsyncreturns up to the settled amount, cumulatively across several refunds of the same settlement (REFUND_EXCEEDS_SETTLEDotherwise). It refunds the overdrawn part first, then the grants the settlement drew from, newest draw first; a grant that has since expired is replaced by a newPromobucket living 30 days (RefundGrace), so a refund is never swallowed by an expiry.
AdjustAsync is the manual correction (positive adds a never-expiring promo bucket, negative consumes
like a settlement), and ExpireAsync zeroes every expired bucket, one Expire entry per grant, keyed
by the grant so a rerun is a no-op.
The reserve → settle → refund lifecycle and the refundable failure classes
The per-step gate is ICreditGate / CreditGate
(inference/src/ReelBolt.WorkflowEngine/Services/Billing/), a thin layer over the ledger that adds a
per-org in-process semaphore, crash-resume, and the failure-class policy. Its options bind the
Billing section, and its enabled flag defaults to false, so a self-hosted ReelBolt never holds a
credit and writes nothing.
ReserveAsyncholds credits for one step. It is idempotent per(stepResultId, attempt)(keygate-reserve:{stepResultId:N}:{attempt}), so a crash-resumed step gets its existing reservation back rather than a second hold. Before touching the ledger it checks the plan entitlement the step needs (Rooms→RoomsAllowed;VideoGeneration→VideoGenAllowedand a non-zero daily USD budget) and refuses withPLAN_NOT_ENTITLED:{requires}. A zero estimate (a BYO-only or runner step) short-circuits to allowed with no reservation and no writes. Any exception is a refusal (GATE_UNAVAILABLE) — a gate that cannot be evaluated must not hand out free compute.SettleAsynccharges the actual and releases the hold in one entry, keyedgate-settle:{reserveId}.ReleaseAsynccloses the hold of a failed step and decides byFailureClass:PlatformProviderError,PlatformInternaland anything unrecognised refund (the switch is a default-refund deny-list, not an allow-list), whileUserInput,UserCancelled,ByoProviderError,InsufficientCreditsandEntitlementDeniedkeep what was actually consumed. If the step had already been charged and then turned out to be our fault, the release refunds the settlement under keygate-refund:{reserveId}, so a retry is never billed twice.ReleaseExecutionAsyncreleases every still-open hold of an execution as a platform failure. It is whatCreditPlatformFailureHookcalls when the execution-recovery sweeper finds an orphaned run, so an engine instance lost mid-run does not cost the customer credits.
FailureClass is the taxonomy that drives this — eight append-only members in the same
Enums.cs: None, UserInput, UserCancelled, ByoProviderError, PlatformProviderError,
PlatformInternal, InsufficientCredits, EntitlementDenied. The design record's rule for the last
of these is deliberate: an unmapped code or unknown exception is PlatformInternal, so it refunds —
charging a user for a failure we cannot explain is the worse error.
Overdraft tolerance exists because a step's real usage cannot be predicted exactly, and failing a
multi-million-token step at 99% of its estimate would be worse than a small overdraft. The gate
reserves at the estimate but passes the ledger a negative floor: headroom = min(estimate x 20%, floorCredits), where the tolerance is Billing:OverdraftTolerancePercent (default 20) and the floor
is Billing:OverdraftFloorCredits (default 100 credits). So a step may finish up to 20% over its
reservation, never more than 100 credits over; the overage is charged, and the balance may go
negative, which blocks further reservations until the organization tops up.
Entitlements and the plan catalog
There is exactly one entitlement contract in ReelBolt: IOrganizationEntitlements
(inference/src/ReelBolt.Shared/Tenancy/), returning an immutable OrganizationEntitlementSet.
Tenancy, billing and compute all read it, and nothing defines a second one. The set carries Tier
(the plan key, or selfhost when billing is off), ByoAllowed, MaxConcurrentExecutions,
VideoGenerationDailyBudgetUsd, MaxOutputHeight, Watermark, RoomsAllowed, VideoGenAllowed,
Priority, StorageGb and RunnerAllowed. "No limit" is a sentinel
(OrganizationEntitlementSet.UnlimitedInt / UnlimitedUsd) rather than null, and 0 means "not
entitled at all" where that reading makes sense.
How an organization's set is chosen (PlanBackedOrganizationEntitlements): with
Billing:Enabled=false every org gets OrganizationEntitlementSet.Unlimited with tier selfhost.
Otherwise the org's entitling subscription's plan is used — Trialing, Active and PastDue all
entitle, so a plan survives the provider's dunning window — and an org with no such subscription falls
back to the plan named by Billing:TrialPlanKey (default trial). A missing trial plan row (an
unseeded catalog) degrades to the configured default rather than throwing. Results are cached per org
for Billing:EntitlementCacheSeconds (default 30), and the webhook processor calls
IEntitlementCache.Invalidate(orgId) on a subscription change so the change applies at once in the
process that handled it; the other service waits out the TTL. Callers must treat the result as a
per-call decision and never cache it beyond a request or a step.
Where entitlements are enforced is PlanGate plus PlanGateService (Inference API) and the
credit gate (engine). The gate returns a machine-readable violation (code, feature, message,
limit, current) with an HTTP status: 402 PLAN_NOT_ENTITLED when the plan does not include a step
type (the three room steps, or VideoGenerate), checked at save time in the workflow create and
update path and again at run time; 429 CONCURRENCY_LIMIT_REACHED when the org is at its
concurrent-execution cap, counted from its Running executions; and 413 STORAGE_QUOTA_EXCEEDED when
an upload would pass the plan's storage quota, measured live as the sum of the org's project-file
sizes rather than from the daily sampler, which lags by up to a day.
The plan catalog is one file, inference/src/ReelBolt.Inference.Api/Billing/plans.catalog.json,
owned by billing and never edited by hand in the database. The Inference API upserts it into the
plans table at startup: unchanged rows are not rewritten, a plan removed from the catalog is
deactivated rather than deleted (subscriptions reference it), the upsert never touches
provider_price_ids_json (which is per-environment data), and an entitlement the set cannot represent
fails startup rather than serving a half-parsed plan. The catalog carries the display fields the
marketing site renders (id, name, kind, priceMonthly, minSeats, includedCredits,
creditsOneTime, highlights, limits, cta, featured) and the billing-only fields
(entitlements, topUpAllowed). The shipped rows are Trial (one-time 100 credits, watermark, 720p,
no BYO, no rooms, no video generation), Creator ($29, 1,000 credits a month, BYO), Pro ($79, 3,500,
rooms and video generation), Team ($49 per seat with a three-seat minimum, 1,500 credits per seat
pooled) and Enterprise (contact only), plus the top-up packs (500, 1,000 and 3,000 credits). Every
entitlement and limit value in it is a placeholder until the pricing work package B13 sets them from
measured costs.
The mirror and the drift check. /site cannot read the canonical file (its Docker build context
is ./site only), so it vendors a byte-identical copy at site/lib/pricing/plans.catalog.json,
loaded and validated by site/lib/pricing/index.ts (schema version, currency, unique ids, known kinds
and CTAs, prices present exactly on priced plans, a minimum seat count on a per-seat plan). CI enforces
the mirror with a diff step that runs for the site app before its typecheck and build:
diff -u inference/src/ReelBolt.Inference.Api/Billing/plans.catalog.json site/lib/pricing/plans.catalog.json. Editing the canonical catalog therefore means copying it to
/site in the same change, or CI fails.
Credits reach an organization through three paths, all of them in the ledger and all of them keyed: the
trial grant (one per normalized email, only for a verified Owner of a Personal workspace, key
trial:{orgId:N}), the monthly plan refill (keyed per subscription period, see below) and a
top-up (keyed per provider transaction).
The IBillingProvider boundary
Everything provider-specific sits behind IBillingProvider
(inference/src/ReelBolt.Inference.Api/Services/Billing/IBillingProvider.cs, decision D11): the
platform sees only neutral shapes, so changing provider is one new class. The interface is Name,
IsConfigured, CreateCheckoutAsync, CreatePortalSessionAsync, VerifyAndParseWebhook,
ChangeSeatsAsync, TopUpPriceId and TopUpPacks. A provider implementation converts its own payload
into a BillingEvent — a neutral record with Provider, EventId, EventType and OccurredAt, one
of four kinds (SubscriptionChanged, TopUpPaid, PaymentRefunded, or Ignored for anything the
platform deliberately does not act on), the organization, customer and subscription ids, the status
and billing period, the priced line items, and the resolved top-up pack size.
BillingEventProcessor never sees a provider payload shape.
A checkout never grants anything: it returns a hosted-checkout URL and a transaction id, and the subscription or the credits change only when a verified webhook is processed. The order of authority is deliberate and is stated in the processor's own documentation — webhooks are authoritative; a checkout redirect or a client call changes nothing.
The shipped implementation is Paddle. Its signature scheme, event types, base URLs, checkout and portal
calls, seat changes, the Billing__Paddle__* configuration keys, the endpoint table and the operator
setup checklist are all in docs/billing.md; this page deliberately does not repeat them.
The webhook pipeline and its idempotency keys
The receiver is POST /api/v1/billing/webhooks/{provider}, anonymous and authenticated by the
provider's signature over the exact raw body alone (see docs/billing.md for Paddle's
header shape and the nginx treatment). Once a payload is verified, BillingEventProcessor.ProcessAsync
applies it with exactly-once effects:
- It checks whether the
(provider, event id)pair is already stored and answersDuplicatewithout doing anything if it is. - Inside one database transaction it inserts the
billing_webhook_eventsrow (provider, event id, event type, the raw payload asjsonb), protected by a unique index on(provider, event_id). A concurrent delivery of the same event loses that index and rolls back as aDuplicate— the winner owns the work. - It applies the effect, marks the row processed and commits. The row, the subscription change and every ledger entry commit together, so a crash leaves nothing half done, and a failure rolls the row back, which is what makes the provider's retry run the event again from scratch.
The outcome is one of four: Applied (state changed, or an event the platform deliberately ignores),
Duplicate, Stale (older than the newest event already applied to that subscription) and
Unresolved (recorded, not actionable — no organization or no known plan, with the reason in the
row's error column). All four are a 200 to the provider; only an exception is a 500, which is
what asks the provider to retry.
Credit effects carry their own idempotency keys as a second line of defence, so a differently numbered event describing the same payment cannot grant twice:
| Effect | Ledger key |
|---|---|
| Monthly plan refill | plan:{provider}:{subscriptionId}:{periodStart UTC} |
| Top-up grant | topup:{provider}:{transactionId} |
| Refund clawback | refund:{provider}:{adjustmentId} |
The refill key is the reason the created, updated and renewal events that all describe one period grant exactly once between them, while a genuine renewal (a new period start) grants again.
The subscription state machine
A subscriptions row is ReelBolt's provider-neutral view of what an organization bought: the plan,
the provider and its subscription id, the status, the billing period, the seat count, whether a
cancellation is scheduled, and last_event_at.
- Statuses are
Trialing,Active,PastDue,PausedandCanceled, mapped from the provider's own status.Trialing,ActiveandPastDueentitle; a scheduled cancel setsCancelAtPeriodEndand leaves the subscription active until the period ends. - Ordering. Providers do not deliver events in order, so every applied change stamps
last_event_atwith the event'soccurred_at, and an event that is not newer is recorded asStaleand not applied. A lateactiveupdate therefore cannot resurrect a subscription that a newer event cancelled. - Resolving the organization goes from the existing subscription row, to the organization id the platform put into the checkout's custom data, to the stored billing customer. An event naming no known organization is stored with a note and acknowledged, not retried forever.
- Plan and seats. The plan is found by matching a subscription item's price id against the plan's provider price map; a per-seat plan takes its seat count from the item quantity and never drops below the plan minimum. A first paid subscription cancels any other non-cancelled subscription of the organization, which is how the internal Trial row is replaced.
- Refill. An
ActiveorTrialingsubscription with a known period start grants the plan's monthly credits (times the seat count for a per-seat plan), expiring at the period end.PastDue,PausedandCanceledgrant nothing, and credits already granted stay until their own expiry; a plan or seat change inside a period does not grant extra credits until the next refill. - Top-up. A completed payment for a configured top-up price grants the pack's credits as a
TopUpbucket expiring 12 months later (BillingEventProcessor.TopUpExpiryMonths). The pack size comes from the price id that was actually paid, never from custom data. - Refund. A full refund of a top-up claws the credits back as a negative ledger adjustment, so the balance may go negative if they were already spent. Only that case is automatic: a partial refund, or a refund of a subscription payment, is recorded with a note and needs a human decision.
What is designed but not built yet
Two of the metering and billing work packages are not finished, and anything that reads as a live per-step charge should be read against this list:
- The credit gate is not wired into step execution (WP B6b).
ICreditGateis registered as a singleton in the WorkflowEngine, andCreditPlatformFailureHookreplaces the recovery sweeper's no-op hook, but nothing in the executor callsReserveAsyncorSettleAsyncper step yet, and no production code path assigns aFailureClass— the only assignment in the source is the platform failure hook'sPlatformInternal. Today the gate therefore releases holds for platform-lost executions and does nothing else: the reserve → settle → refund lifecycle described above is implemented and unit-tested at the ledger and gate level, but a step does not open a reservation. VideoGeneratestill charges the legacy USD path (WP B6c). Its budget gate (VideoGenerationBudgetGate), theexternal_generation_jobsspend ledger withcost_estimate_usdandactual_cost_usd, and theBUDGET_EXCEEDED/MaxSpendUsdcaps are still what stops a run overspending; theVideoGenSecondsusage row is written alongside them (keyedvideogen:{jobId:N}) as the credit-side record, so the two never double count, but only the USD path is enforced today. Also unmetered on the failure path: computed seconds are recorded on success only, so a failed render or encode is not yet charged before being refunded by class.