Video Generation
Generate new video clips from a text prompt inside a workflow. You register a MiniMax provider
row, add a StepType.VideoGenerate step with a prompt or a plan plus a spend cap, and the step
submits one clip per take, polls it, downloads it, normalizes it with ffmpeg, and stores it in the
project's object storage. Every submission passes a pre-submit budget gate first.
Issue #93 phase 1 built the provider layer:
one provider (MiniMax), one operator-written prompt per step. Phase 2 added the planner. A
ShotDirector agent reads the same bounded VideoAnalyze view the story editor decides from and
plans the shots; the same deterministic step then buys one clip per planned shot under the same
spend caps. A VideoGenerate step therefore has two plan sources: Inline, phase 1's single
prompt, and StepRef, a planner step's plan. For the surrounding engine (step types, executors,
provider resolution) see CLAUDE.md; for the editing steps that consume a generated clip,
including where the compile step places one, see video-editing.md.
Table of Contents
- How a generated clip reaches the rest of a workflow
- Planning the shots
- MiniMax facts this feature relies on
- Set up a MiniMax provider
- Step config reference
- The budget gate
- Egress: sending the project's own frames to a provider
- Job lifecycle and crash-resume
- Step output
- How a planned clip is addressed
- Where the clips go in the edit
- Accepted risks
- Step cache
- Storage layout and normalization
- Configuration
- Workflow templates
- Out of scope for phases 3, 4 and 5
How a generated clip reaches the rest of a workflow
VideoGenerateStepExecutor sets the step result's OutputStorageKey to the storage key of the
first take that was delivered. A later VideoAnalyze step reads it through its
VideoSourceRef like any other rendered output:
Source: PreviousStepOutputpicks the latest step in the execution with a non-nullOutputStorageKey.Source: StepOutputwithStepOrderset to theVideoGeneratestep picks it explicitly.
See Picking the source video. Clips after the first —
further takes on the Inline path, further planned clips on the StepRef path — are stored and
listed in the step's clips output, but no VideoSourceRef kind addresses them. On the plan path
the manifest is what maps a clip id to its object (see
How a planned clip is addressed).
VideoGenerate is a deterministic, non-LLM step: it has no agent of its own and reuses the seeded
AgentType.VideoTransform row to satisfy WorkflowStep.AgentDefinitionId, like
VideoAnalyze/VideoCompile. It never throws; every failure is a structured JSON output.
Planning the shots
AgentType.ShotDirector is an LLM agent that plans generated shots from the same bounded
VideoAnalyze view the story editor decides from. It is an ordinary Agent step — there is no new
step type for it — implemented in
inference/src/ReelBolt.WorkflowEngine/Agents/Production/ShotDirectorAgent.cs, with
OutputSchemaName = "GeneratedShotPlanOutput". It is seeded as a built-in agent row; its system
prompt exists in two places, the class's own DefaultPrompt and DatabaseSeeder.BuiltInAgents,
and ShotDirectorPromptConsistencyTests asserts the two stay byte-identical.
It reads two things, which is why a template runs it with AgentInputContextMode.FullWorkflow:
- the bounded
VideoAnalyzeview, whoseviewoffers thes{n}(shot),g{n}(silence gap) andt{n}(transcript segment) cut-anchor ids a planned shot may hang off; - the story editor's decision, so it can avoid anchoring a shot to a moment the edit cut away.
It emits GeneratedShotPlanOutput — { shots, planRationale }, where each shot is:
{
"purpose": "Cutaway",
"anchor": "s3",
"firstFrame": null,
"lastFrame": null,
"camera": "PushIn",
"duration": "Short",
"prompt": "A slow push across a sunlit desk…",
"reason": "Covers the weakest stretch of picture inside a kept run."
}
ShotPlan is exactly those eight properties and every one of them is a plain string
(firstFrame and lastFrame are nullable). purpose is one of Cutaway, SeamBridge, Extend,
ColdOpen, EndCard or ScreenContent; duration is one of Short, Medium or Long. An
empty shots list is a fully valid outcome — an edit needs no generated shot.
The no-number invariant
No property anywhere in ShotPlan or GeneratedShotPlanOutput is a numeric or time-bearing type.
GeneratedShotPlanOutputInvariantTests is a reflection test that fails if one is added, fails if
any ShotPlan property stops being a string, and pins both exact property sets. The planner is
therefore structurally incapable of emitting a timestamp, a frame index, a duration in seconds, a
resolution, a pixel coordinate or a URL.
The reason matters more than the test. The prompt is the only value the model authors that leaves
the building; every other value it emits is either an opaque id that the analysis artifact offered
it, or an enum word. First-party C# turns those words into provider parameters:
ShotPlanParameterMapper appends an English camera clause to the prompt and snaps the duration
word to a number of seconds, while ShotPlanCostEstimator and VideoGenerationPricing do the
money. So nothing the model writes is interpreted as a number, a time or a provider parameter.
ShotPlanParameterMapper.CameraWords is the canonical 11-word camera vocabulary. SnapDurationSeconds
maps Short to the model's minimum duration, Long to its maximum and Medium to the midpoint,
clamped to the model's range in every case. An unknown duration word throws rather than
defaulting, because the snapped duration drives the spend estimate and the ledger row; an unknown
camera word is simply omitted from the prompt, which is a documented degradation because nothing
downstream depends on it.
The prompt sanitizer
GeneratedPromptSanitizer (ReelBolt.Shared/Inference/GeneratedPromptSanitizer.cs) is the one
shared prompt-hygiene pass for model-authored prose; phase 1's private
VideoGenerateStepExecutor.SanitizePrompt and its duplicate length constant are gone.
DefaultMaxPromptLength is 7000, and MaxPromptLengthFor(InferenceProviderKind) is total over
every provider kind (MiniMax 7000, default 7000) so a future kind cannot silently receive a zero
cap and disable itself. StripControlCharacters drops every control character except \n, \r
and \t.
SanitizedPrompt carries Text, MaxLength, IsEmpty and ExceedsMaxLength, and it does not
truncate. An empty or over-length prompt fails closed with INVALID_PROMPT instead, so a provider
is never sent a prompt the model did not write. On the Inline path that fails the step; on the
StepRef path it degrades that one clip and the step continues.
The planner's tool scope, and the money boundary
ToolGroupCatalog grants AgentType.ShotDirector exactly ToolGroup.ProjectRead plus
ToolGroup.WorkflowControl — the same minimal read-only scope as VideoStoryEditor,
MusicSupervisor and SoundDesigner. There is no sandbox tool, no render tool and no project-write
tool, and ToolScopingDriftGuardTests.ShotDirector_is_granted_no_sandbox_render_or_spending_tool
asserts the grant rather than leaving it to inspection.
The conclusion this is built for: the planner cannot spend money. The planning step is an
ordinary agent call, cheap and bounded, and it has nothing that can reach a provider. Only the
deterministic VideoGenerate step buys anything, and only within its own config — its MaxSpendUsd
and the budget gate below.
MiniMax facts this feature relies on
Checked against MiniMax's own documentation on 2026-09-28. Each fact links the page that states it.
| Fact | Value | Source |
|---|---|---|
| Base URL | https://api.minimax.io | create |
| Auth | Authorization: Bearer <api key> | create |
| Create task | POST /v2/video_generation, body model, content, resolution, duration, ratio; returns task_id | create |
| Text-to-video input | content holds a single {"type": "text", "text": ...} item | create |
| Prompt length | at most 7000 characters per text item | create |
MiniMax-H3 | 768P or 2K; duration integer 4–15 s | create |
MiniMax-H3-Max | 480P or 768P (no 2K); duration integer 5–15 s | create |
| Ratio for text-to-video | required, one of 21:9, 16:9, 4:3, 1:1, 3:4, 9:16; adaptive is not allowed | create |
| Prices (pay-as-you-go) | MiniMax-H3: 768P $0.08/s, 2K $0.13/s. MiniMax-H3-Max: 480P $0.05/s, 768P $0.08/s | pricing |
| Query task | GET /v2/query/video_generation/{task_id} (task_id is a path parameter) | query |
| Task statuses | queued, running, succeeded, failed, cancelled | query |
| Download | on succeeded, task.content.url is a time-limited download URL; query again for a fresh one | query |
| Query window | only tasks from the last 7 days are queryable | query |
| Metered usage | task.usage.total_seconds, "input seconds + output seconds", returned only on success | query |
| Poll cadence | 10 seconds recommended | guide |
| Errors | 400 bad_request, 401 authorized, 402 insufficient_balance, 422 unprocessable_entity, 429 rate_limit, 500 server | create |
| Idempotency | the create endpoint documents no idempotency key | create |
The per-model resolutions, durations, ratios and prices above are hard-coded in
MiniMaxVideoGenerationClient.Describe (ReelBolt.Shared/Inference/MiniMaxVideoGenerationClient.cs).
If MiniMax changes them, that method is the one place to update.
Set up a MiniMax provider
Admin → Inference Providers → New, or POST /api/v1/inference-providers:
| Field | Value |
|---|---|
| Kind | MiniMax |
| Capability | VideoGeneration |
| Endpoint | https://api.minimax.io |
| API key | your MiniMax API key |
| Timeout | per-HTTP-request timeout in seconds, applied to submit, poll and download calls |
| IsDefault | set it to make this the row a step uses when its ProviderId is unset |
The step's Model field picks the model; the executor does not read the provider row's model name.
The admin form prefills the endpoint as https://api.minimax.io, the base URL MiniMax documents
for the v2 API, when you pick MiniMax.
The API validates the kind/capability pair on create, update and test and returns 400 for:
Capability: VideoGenerationwith any kind other thanMiniMax.Kind: MiniMaxwith any capability other thanVideoGeneration.
There is no env: sentinel for MiniMax. The stored key is decrypted and sent verbatim as the
bearer token.
Test a provider without spending
POST /api/v1/inference-providers/{id}/test (or /test for an unsaved config) never creates a
task, so it costs nothing. MiniMaxVideoGenerationClient.ProbeAsync sends one
GET /v2/query/video_generation/{random GUID} with the configured key, bounded by a 15-second
timeout (the provider row's Timeout, if set lower, cuts it shorter), and classifies the answer:
| Answer | Result |
|---|---|
401 or 403, or a MiniMax error body whose error.type is authorized_error (any status) | ok: false, error "Provider rejected authentication." |
500 with error.type server_error and a message starting "record not found" — what MiniMax actually answers for a task id that does not exist | ok: true |
A network error, the timeout, 408, 429 or any other 5xx | ok: false, error "Provider test failed: …" naming what happened |
| A non-JSON body, or JSON that is neither MiniMax envelope below | ok: false, same error form |
Any other status whose body is {"task": {...}} | ok: true |
Any other status whose body is MiniMax's error envelope {"type": "error", "error": {"type": ..., "message": ...}} with an error.type other than authorized_error (for example bad_request_error, "invalid task_id") | ok: true |
The result's responsePreview (on ok: true and on an auth rejection) carries the status and at
most a short, key-scrubbed vendor error.type/error.message, never the raw body. On the saved
provider path the outcome is persisted to LastTestAt/LastTestOk/LastTestError.
A passing test means the endpoint answered with MiniMax's query envelope and did not reject the
key. MiniMax does not document what the query endpoint returns for a task id that does not exist,
but observed behaviour (2026-10-04) is that it checks the key first — an invalid key gets 401
authorized_error — and then answers a valid key's lookup of an unknown id with HTTP 500,
server_error, "record not found (1000)". That 500 is why the probe special-cases it: treating
every 5xx as a failure made the test fail for every working key. Any other 5xx, including a
server_error with a different message, still fails.
Step config reference
WorkflowStep.VideoGenerateConfigJson (column video_generate_config_json, jsonb) holds a
JSON-serialized VideoGenerateStepConfig (ReelBolt.Shared/Workflows/VideoGenerateStepConfig.cs).
Property names are camelCase in JSON; enums are strings.
| Property | Type | Default | Notes |
|---|---|---|---|
PlanSource | VideoGeneratePlanSource | Inline | Inline generates Takes clips from InlinePrompt; StepRef generates one clip per planned shot. StepRef is appended after Inline so stored payloads keep their meaning. |
InlinePrompt | string? | null | The prompt. Required on the Inline path, see below. Ignored on the StepRef path. |
Plan | ExtractInputRef? | null | The planner step's GeneratedShotPlanOutput. Required when PlanSource is StepRef, see below. |
Analysis | ExtractInputRef? | null | The VideoAnalyze step whose artifact offered the firstFrame/lastFrame ids. Required when any planned shot names one, see below. |
Model | string | "MiniMax-H3" | MiniMax-H3 or MiniMax-H3-Max. Any other value fails CONFIG_INVALID. |
Resolution | string | "768P" | Must be one of the model's resolutions, matched exactly (case-sensitive, no trimming). |
DurationSeconds | int | 6 | Seconds per clip on the Inline path. On the StepRef path each shot's own duration word is snapped instead, so this value does not drive the clips — but it is still validated against the model's range, and it is still part of the ledger's request_hash on both paths. |
Ratio | string | "16:9" | Must be one of the model's ratios, matched exactly. |
ProviderId | Guid? | null | A VideoGeneration provider row. If unset, disabled, deleted or of another capability, the default VideoGeneration row is used. |
Takes | int | 1 | Inline: clips to generate, clip-1..clip-N. StepRef: takes per planned shot, and the clip ids become v{n}t{k}. |
MaxClips | int | 1 | Author-set hard ceiling on clips for the whole step, validated independently of Takes. |
MaxSpendUsd | decimal | 0 | Step spend cap in USD. No usable default: 0 fails validation. |
AllowSourceMediaEgress | bool | false | Consent to send frames of the project's own footage to the provider as keyframes. See Egress. |
PollIntervalSeconds | int | 10 | Delay between polls, 1 to 300 inclusive. MiniMax recommends 10. |
TimeoutSeconds | int | 1800 | Per-take polling budget, 60 to 7200 inclusive and at least PollIntervalSeconds, counted from when the take holds a provider slot. A take still unfinished when it elapses ends POLL_TIMEOUT. |
Validate() returns every rule that fails; at run time any error ends the step CONFIG_INVALID
before the budget gate reserves anything, so nothing is submitted or charged. A missing
VideoGenerateConfigJson, invalid JSON or a JSON null also ends the step CONFIG_INVALID.
MaxSpendUsdmust be greater than 0.Takesmust be between 1 and 4 inclusive.MaxClipsmust be between 1 and 4 inclusive.- If
MaxClipsis in range,Takesmust not exceedMaxClips. DurationSecondsmust be greater than 0.PollIntervalSecondsmust be between 1 and 300 inclusive.TimeoutSecondsmust be between 60 and 7200 inclusive.- If both are in range,
TimeoutSecondsmust be at leastPollIntervalSeconds; otherwise the take would be paid for and time out before its first poll. Planis required whenPlanSourceisStepRef.PlanandAnalysis, when set, must be sourcedFrom = PreviousorFrom = Step, andFrom = Steprequires aStepOrder.
MaxClips is a ceiling counted across the whole step, spent in plan order — shot major, take
minor. On the StepRef path every clip past it is reported skipped with
skipReason: MAX_CLIPS_EXCEEDED and consumes no ledger row, and its bound is still 1..4.
Validate(capabilities) runs after the provider resolves, against Describe(Model), and adds:
DurationSecondsmust lie within the model's minimum and maximum.Resolutionmust be a key of the model's price table (ordinal match).Ratiomust be one of the model's allowed ratios (ordinal match).
It deliberately adds no StepRef rule: DurationSeconds stays the Inline path's per-clip
duration, while the plan path snaps each shot's own duration word, so it is in range by
construction.
On the Inline path the executor resolves the provider first — the prompt's length cap is per
provider kind — and then sanitizes the prompt:
- An empty or whitespace-only
InlinePromptfailsINVALID_PROMPT. - A prompt over the resolved kind's cap (7000 characters for MiniMax) fails
INVALID_PROMPT; it is not truncated. - Control characters other than
\n,\rand\tare stripped first, so a prompt made only of control characters is rejected as empty.
On the StepRef path each planned shot's prompt goes through the same pass, but an empty or
over-length prompt fails only that clip.
The Analysis reference is not required by Validate(), because the config alone cannot know what
the plan will contain. The executor enforces it: a plan in which any shot names firstFrame or
lastFrame and no Analysis is configured ends the step CONFIG_INVALID, naming the offending
shots, before anything is reserved.
Save-time validation
The Inference API rejects a VideoGenerate step whose config would fail Validate() before it is
stored, using StepConfigSaveValidator with the executor's own JSON options. A missing or blank
videoGenerateConfigJson, invalid JSON and a JSON null are rejected too. Each error reads
Step <order> (VideoGenerate): <rule>.
| Path | Rejection |
|---|---|
POST and PUT /api/v1/projects/{projectId}/workflows[/{id}] | 400 {"message": "<errors, one per line>", "errors": [...]}, nothing written |
POST /api/v1/projects/{projectId}/workflows/assistant/apply (assistant Apply) | same 400 {message, errors}, before the create or update branch writes anything |
Assistant ApplyWorkflow tool | tool result {"error": "<errors, one per line>", "errors": [...]}, nothing written |
Assistant ValidateWorkflow tool and POST .../workflows/assistant/validate | a result with valid false and the same errors, so validation never passes a step Apply would reject |
Save time checks only the provider-free rules — which now include the structural Plan and
Analysis rules, so a plan-sourced config with no Plan is rejected before it is stored. The
per-model duration, resolution and ratio checks and the prompt checks need the resolved provider or
run in the executor, so they still fail the step at run time.
The web workflow builder (web/lib/utils/video-generate-validation.ts) blocks Save client-side
and shows Step N: <first error>. It applies the Validate() rules it can check without a resolved
provider, and adds the prompt rules, which the server checks only at run time:
InlinePromptmust be non-blank and at most 7000 characters.InlinePromptmust contain no control characters other than\n,\rand\t. The executor strips those at run time instead of failing.- A
VideoGeneratestep with no config is reported as "Max Spend (USD) must be greater than 0", the first error its default config would produce.
Example Inline step config:
{
"planSource": "Inline",
"inlinePrompt": "A slow dolly shot across a sunlit desk with a laptop showing a dashboard",
"model": "MiniMax-H3",
"resolution": "768P",
"durationSeconds": 6,
"ratio": "16:9",
"takes": 2,
"maxClips": 2,
"maxSpendUsd": 1.0
}
That step estimates 2 × 6 s × $0.08/s = $0.96 and passes the $1.00 step cap.
Example plan-sourced step config:
{
"planSource": "StepRef",
"plan": { "from": "Step", "stepOrder": 3 },
"analysis": { "from": "Step", "stepOrder": 1 },
"model": "MiniMax-H3",
"resolution": "768P",
"durationSeconds": 6,
"ratio": "16:9",
"takes": 1,
"maxClips": 4,
"maxSpendUsd": 5.0
}
plan names the ShotDirector step and analysis the VideoAnalyze step whose offered ids
resolve any keyframe a shot asks for. durationSeconds is not read for pricing on this path — each
shot's own duration word is snapped instead — but it still has to be a duration the model offers,
because Validate(capabilities) checks it unconditionally.
The budget gate
IVideoGenerationBudgetGate.ReserveAsync (WorkflowEngine/Services/VideoGeneration/) checks
every cap and writes every take's ledger row as one atomic operation, before anything is
submitted.
Pricing is fail-closed. The executor prices each take with
VideoGenerationPricing.EstimateClipCostUsd: the model's per-second price for Resolution times
DurationSeconds. An unknown model, an unpriced resolution, or a non-positive estimate throws,
and the step fails CONFIG_INVALID. An estimate is never 0.
The reservation, in order:
- No takes, a take estimate ≤ 0,
MaxSpendUsd≤ 0, or aGuid.Emptystep-result id → denyCONFIG_INVALID. - A take whose
(ExecutionId, RequestHash)already has aReservedorSubmittedrow is returned as existing: not inserted again and not charged again. - Step cap: the sum of every take's estimate (new and existing) >
MaxSpendUsd→ deny. - Project daily cap: the project's charged spend today + the new takes' estimate >
VideoGeneration:DailyBudgetUsd→ deny. A cap ≤ 0 denies everything. - Global daily cap: charged spend today across all projects + the new takes' estimate >
VideoGeneration:GlobalDailyBudgetUsd→ deny. A cap ≤ 0 denies everything. This stops extra projects from multiplying the per-project cap. - Otherwise one
Reservedrow is inserted per new take and every take is returned in order.
All or nothing. A denial writes no rows and submits nothing; the step fails
BUDGET_EXCEEDED (or CONFIG_INVALID) with the gate's reason in meta.error. There is no
partial fulfilment.
What counts as spent. "Today" is rows with CreatedAt on or after UTC midnight. Every row
except Failed and Cancelled is charged — Reserved, Submitted, Running and Succeeded.
A charged row costs max(CostEstimateUsd, ActualCostUsd, 0), so a vendor-reported actual below
the estimate never frees headroom, and no row can credit the ledger. Failed and Cancelled
rows are released.
Concurrency. The read-then-insert runs under a process-wide lock, inside a transaction, and
on PostgreSQL under pg_advisory_xact_lock with a single global key. Two concurrent steps,
in the same or different projects and engine processes, cannot both pass a cap only one of them
fits under.
Pricing a whole plan
On the StepRef path the executor prices the whole plan before reserving anything.
ShotPlanCostEstimator.EstimateWholePlanUsd sums, over every planned shot,
Describe(model).PricePerSecond × snapped duration × takes. Every per-clip price comes from
VideoGenerationPricing, the one pricing path, so the estimator cannot drift from the ledger.
It fails closed the same way: an unpriceable resolution or an unresolvable duration word throws and
the step ends CONFIG_INVALID rather than contributing a silent 0, because a zero estimate passes
every cap. An empty plan is the one legitimate zero — there is genuinely nothing to buy.
The whole-plan estimate is compared against MaxSpendUsd before the gate's reservation.
Exceeding it ends the step BUDGET_EXCEEDED with zero submits and zero ledger rows.
A plan larger than MaxClips still has its whole-plan estimate gated. The cap binds the plan
the author wrote, not the trimmed subset that would actually be dispatched. The gate is handed only
the trimmed take list, so on its own it could never see that the full plan was over budget — which
is exactly why this check sits outside it. The gate call that follows
(IVideoGenerationBudgetGate.ReserveAsync) then covers the effective, post-trim plan exactly as in
phase 1, and meta.planEstimatedUsd is that effective number.
Egress: sending the project's own frames to a provider
A keyframe is not metadata about the user's footage — it is the user's footage, and once it has
been posted to a provider it cannot be recalled. AllowSourceMediaEgress (default false) is the
workflow author's explicit consent to send it, and GeneratedShotEgressPolicy decides it per shot.
The executor evaluates that policy against every shot in the plan first. A single refusal fails
the whole step with code EGRESS_REFUSED and the policy's reason, with zero submits and zero ledger
rows. A partially generated plan is a half-edit the operator cannot use, so the refusal has to be
actionable — grant consent, or edit the plan — rather than silently dropping one shot.
With consent granted the step proceeds, but the frame is still not transported. In this wave
VideoGenerationRequest has no image field, so a consented keyframe-requiring shot is submitted
best-effort from its prompt alone. Each such clip is recorded rather than silently ignored, in the
per-clip manifest:
| Field | Value |
|---|---|
keyframeRequested | true |
keyframeApplied | false |
keyframeReason | keyframe_transport_not_available |
meta.keyframesRequested and meta.keyframesNotApplied count them. Transporting the frame is wave
B's work — Higgsfield's presigned upload is where it lands — and it is why these fields exist now.
When a shot names a frame at all, the config's Analysis reference is required, and the named id
must be one the analysis artifact actually offered the planner. An id that was never offered records
keyframeReason: "keyframe_unresolved" on that clip, which is likewise not fatal.
Job lifecycle and crash-resume
Each take has one row in external_generation_jobs (WorkflowEngine-owned, no foreign keys, so a
row outlives its execution, step result and project). Columns: id, execution_id,
step_result_id, project_id, clip_id, request_hash, provider_id, remote_job_id,
status, cost_estimate_usd, actual_cost_usd, storage_key, created_at, updated_at.
status is one of Reserved, Submitted, Running, Succeeded, Failed, Cancelled; the
executor never writes Running. step_result_id is the id of the step result for the attempt
that reserved the row. A retry that reuses the row (see Attempts and resume)
keeps the original attempt's id, so the column does not point at the attempt that last polled.
The gate refuses a reservation whose step-result id is Guid.Empty with CONFIG_INVALID.
request_hash is SHA-256 over model, resolution, ratio, duration, take index and the sanitized
prompt.
Status per exit path
Takes run one after another. Every exit writes UpdatedAt and saves with
CancellationToken.None, so a stopped step still records its outcome. failureReason in the
clip output is the code; detail carries the human explanation.
| Exit path | Row status afterwards | failureReason | Charged? |
|---|---|---|---|
| Delivered | Succeeded, StorageKey set | — | yes |
Vendor reports failed | Failed | PROVIDER_INSUFFICIENT_BALANCE if the reason contains 402 or insufficient, else PROVIDER_FAILED | released |
Vendor reports cancelled | Cancelled | PROVIDER_CANCELLED | released |
| Submit returns 4xx other than 408 | Failed | PROVIDER_INSUFFICIENT_BALANCE (402), PROVIDER_AUTH_FAILED (401/403), else PROVIDER_REJECTED | released |
| Submit fails otherwise (5xx, 408, network, malformed 2xx, cancellation) or returns no task id | Reserved, unchanged | SUBMIT_OUTCOME_UNKNOWN | yes |
| Poll throws, or the step is stopped while a submitted take waits or polls | Submitted, unchanged | POLL_ERROR | yes |
TimeoutSeconds elapses | Submitted, unchanged | POLL_TIMEOUT | yes |
Vendor succeeded with no download URL | Submitted, ActualCostUsd set | NO_DOWNLOAD_URL | yes |
Vendor succeeded, then download, normalize or upload throws | Submitted, ActualCostUsd set | DELIVERY_FAILED | yes |
| Fresh take never submitted because the step was stopped | Cancelled | NOT_SUBMITTED | released |
Take reserved by an earlier attempt with no RemoteJobId | Reserved, untouched | RESERVATION_ORPHANED | yes |
| Any other exception | Cancelled if the take was fresh and never submitted; otherwise unchanged | UNEXPECTED_ERROR | released / yes |
Only a vendor-confirmed failure or cancellation, a definite 4xx submit rejection, or a take that never reached the vendor is released. Everything that might have been billed stays charged.
Actual cost
When the vendor reports succeeded, before downloading, the executor records
ActualCostUsd = usage.total_seconds × the model's per-second price for Resolution. A usage of
0 records 0; a missing, null, non-integer or negative usage records null. The executor never
copies the estimate into ActualCostUsd.
Attempts and resume
VideoGenerate is retried like an Agent step: up to WorkflowHardening:MaxStepRetries
attempts (default 3), each calling ExecuteAsync again with the same ExecutionId. On the
dispatch path the step is Failed only when every dispatched take failed, so that is the only
take-driven retry. Every pre-gate refusal is Failed too (CONFIG_INVALID, INVALID_PROMPT,
PROVIDER_NOT_FOUND, EGRESS_REFUSED, BUDGET_EXCEEDED, UNEXPECTED_ERROR) and is retried the
same way, to the same limit: VideoGenerate is deliberately not one of the deterministic step
types capped at a single attempt. Each attempt goes through the budget gate again, within every
cap.
On a retry, the gate hands back the earlier attempt's Reserved/Submitted rows as existing:
- Row has a
RemoteJobId: the take is polled only, never re-submitted. This is how aPOLL_TIMEOUTorPOLL_ERRORtake gets another chance, and how aDELIVERY_FAILEDtake gets a fresh download URL. - Row has no
RemoteJobId: the earlier attempt may have reached MiniMax without recording the task id. The take endsRESERVATION_ORPHANEDwith zero submit calls and the row untouched. Failed/Cancelledrows are not returned as existing, so those takes get a new row and a new submission.
Handling an orphaned reservation
- Read the clip's
detail: it names the ledger row id and itsCreatedAt(UTC). - In the MiniMax console, look for a video task created near that time.
- If a task exists, the money was spent and the clip is not in ReelBolt. If none exists, nothing was billed.
- Either way the reservation stays charged against the project and global caps until the UTC day rolls over. Do not edit the row to free budget unless you have confirmed no task exists.
Per-provider concurrency
VideoGeneration:MaxConcurrentPerProvider (default 2, minimum 1) bounds how many takes hold a
slot on one provider at once, keyed by provider id. A take acquires the slot before it submits
(or, when resuming, before its first poll) and releases it when polling ends, before download
and normalization, so a slow download never blocks another submit. The TimeoutSeconds clock
starts only once the slot is held. The limiter is per engine process.
MiniMax documents no concurrency limit for the v2 H3 API; the default of 2 is ReelBolt's own conservative choice.
Step output
A step that reached the gate emits:
{
"clips": [
{ "clipId": "clip-1", "status": "succeeded", "stored": true, "costUsd": 0.48, "costBasis": "actual" },
{ "clipId": "clip-2", "status": "failed", "failureReason": "POLL_TIMEOUT", "detail": "…", "costUsd": 0.48, "costBasis": "estimate" }
],
"meta": {
"totalCostUsd": 0.96,
"actualCostUsd": 0.48,
"unconfirmedCostUsd": 0.48,
"estimatedCostUsd": 0.96,
"clipsGenerated": 1,
"clipsFailed": 1,
"clipsSkipped": 0
}
}
Both plan sources share this projection, so the Inline path's clip-{n} clips and the plan path's
v{n} clips are shaped identically. A clip carries {clipId, status, failureReason?, detail?, skipReason?, stored, costUsd, costBasis}.
stored is a boolean, not a path — OutputJson reaches later agents under
AgentInputContextMode.FullWorkflow, so no storage key may appear in it (see
How a planned clip is addressed). shotIndex, take,
purpose, anchor and durationSeconds are manifest-only and do not appear here.
costBasis labels each clip's costUsd:
costBasis | When | costUsd |
|---|---|---|
actual | ActualCostUsd is known | ActualCostUsd |
estimate | row charged, actual unknown | the row's CostEstimateUsd, never 0 |
released | row Failed or Cancelled | 0 |
meta.totalCostUsd sums every clip's costUsd; actualCostUsd and unconfirmedCostUsd sum the
actual and estimate clips; estimatedCostUsd sums every take's pre-submit estimate;
clipsGenerated, clipsFailed and clipsSkipped count clips by status.
A plan-sourced step adds:
| Key | Meaning |
|---|---|
planEstimatedUsd | The effective, post-trim plan's estimate — the exact sum of the reservations handed to the gate. |
shotsPlanned | How many shots the planner offered, before the MaxClips trim. |
keyframesRequested | Clips whose planned shot named firstFrame or lastFrame. |
keyframesNotApplied | Of those, the ones whose keyframe was not applied. |
note | EMPTY_PLAN when the planner legitimately offered no shots at all. |
If every take failed, the step status is Failed, meta.failed is ALL_TAKES_FAILED, and
ErrorDetails lists clipId:code pairs, e.g. clip-1:POLL_TIMEOUT, clip-2:PROVIDER_FAILED. With
at least one delivered take the step is Completed.
A step that fails before the gate reserves anything (CONFIG_INVALID, INVALID_PROMPT,
PROVIDER_NOT_FOUND, BUDGET_EXCEEDED, EGRESS_REFUSED, UNEXPECTED_ERROR) emits
{"clips": [], "meta": {"failed": "<code>", "error": "<message>"}}.
How a planned clip is addressed
OutputStorageKey still holds only the first succeeded clip, which is what the UI player and a
later VideoAnalyze step read. Every other planned clip is addressed through the per-clip manifest:
projects/{projectId}/agentFiles/video-analysis/{executionId}/step-{order}-generated-clips.json
VideoGenerateStepExecutor writes it and points WorkflowStepResult.ArtifactStorageKey at it. It
carries clips[].{clipId, shotIndex, take, purpose, anchor, firstFrame, lastFrame, camera, duration, durationSeconds, status, storageKey, costUsd, costBasis, keyframeRequested, keyframeApplied, keyframeReason, failureReason, detail, skipReason} plus the step's meta, and the compile step
resolves generatedClips → that step's ArtifactStorageKey → the manifest.
Clip ids are v{n} in 1-based plan order when Takes is 1, and v{n}t{k} otherwise, so a later
take-selection pass has distinct ids to choose between. The Inline path keeps phase 1's clip-{n}.
Why the manifest exists at all. The step's output JSON no longer carries clips[].storageKey.
It carries clips[].stored instead, so that no storage path is ever model-visible: the output
reaches later agents under AgentInputContextMode.FullWorkflow, and a test asserts the output
contains no storageKey and none of projects/, agentFiles/ or outputFiles/, while the
manifest behind ArtifactStorageKey carries the real keys.
Because the manifest sits under the video-analysis/ prefix and not outputFiles/, the existing
artifact endpoint serves it with no backend change:
GET /api/v1/projects/{projectId}/step-results/{stepResultId}/artifact
It already validates that prefix, so the manifest reads through it exactly like a VideoAnalyze or
VideoCompile artifact does.
Where the clips go in the edit
The compile side gains three appended fields on VideoCompileStepConfig:
generatedClips (ExtractInputRef?, default null), enableGeneratedClips (bool, default
false) and maxGeneratedClips (int, default 8). Placement derives from each planned shot's own
purpose — there is deliberately no per-purpose placement config, because a second config surface
would only let the two disagree.
purpose | What the compile step does |
|---|---|
Cutaway | Replaces picture only over the anchor's window on the OUTPUT timeline. The base audio chain is untouched, so dialogue runs continuously across the cutaway — a J/L-cut. |
ColdOpen | Prepended as its own concat segment, pairing with ProgramFadeIn. |
EndCard | Appended as its own concat segment, pairing with ProgramFadeOut. |
SeamBridge, Extend, ScreenContent | Accepted vocabulary that degrades explicitly to not_supported_in_this_wave rather than being silently misplaced as one of the three above. |
A cutaway's window is [anchorStart, min(anchorStart + clipDurationSeconds, anchorEnd)), on the
output timeline. Every placement failure degrades per clip and never fails the compile; the one
hard config error is enableGeneratedClips with Mode: StreamCopy, which ends the step
GENERATED_CLIPS_REQUIRE_REENCODE because a picture replacement and a prepended segment both need a
filtergraph.
EnableGeneratedClips: false — the default — is byte-identical to the pre-wave compile path: no
extra ffmpeg input, no extra filter, and no new EDL or output key. That backward-compatibility
guarantee is load-bearing, the same one the graphics, music, inserts, grade and SFX blocks carry.
For the filtergraph-level detail — how the replacement is gated, how the concat segments are built,
and what the EDL's generatedClips node reports — see
video-editing.md.
Accepted risks
- No vendor idempotency key. MiniMax's create endpoint has none, so exactly-once submission
is not achievable client-side. If a submit reaches MiniMax but its task id is never recorded
(
SUBMIT_OUTCOME_UNKNOWN, or a crash mid-submit), the take is charged at its estimate and never re-submitted, but ReelBolt cannot track or deliver that clip. The operator reconciles through the console (see Handling an orphaned reservation). - False orphans err on the charged side. A take reserved but not yet submitted when the
process died (for example while waiting for a provider slot) looks the same as one whose submit
reached MiniMax. If the same execution runs the step again it is reported
RESERVATION_ORPHANEDand stays charged, even though nothing was billed. - A cache hit validates only the first take. See Step cache.
Step cache
StepCachePolicy treats VideoGenerate as cacheable. The cache key includes
VideoGenerateConfigJson, so the prompt, model, resolution, duration, ratio, takes and provider
id (when set) are all part of it. A repeated execution with identical inputs replays the prior Completed
result, including a partial one, instead of paying again. A Failed result is never cached.
A plan-sourced step's key also covers the planner's output, transitively: the key is taken over the fully resolved step input, which for a deterministic step type is the preceding step's accumulated output. A plan that changed is therefore a cache miss, and an unchanged one replays without paying twice.
Before serving a hit, the cache HEADs the result's OutputStorageKey, which is only the first
delivered take — on the plan path that means only clips[0]'s object. Objects for other clips are
not checked, so a hit can reference a deleted clip object.
Setting the step's CacheMode to Never only skips the step cache: identical clips are still reused from the spend ledger (next section). To force fresh, paid takes, run with Regenerate (regenerateVideoClips: true).
Clip reuse on re-run (ledger-level)
The step cache only hits when the whole resolved input is byte-identical, so a ReviewLoop
loop-back (the upstream output changed) or a retry that follows a partial failure would still pay
again. Independently of the cache, VideoGenerateStepExecutor therefore reuses clips at the
ledger level. Before the budget gate is called, each take's request_hash (model, resolution,
ratio, duration, take index, sanitized prompt) is looked up in external_generation_jobs for the
same project and provider: a Succeeded row with a StorageKey whose object still exists (HEAD via
ArtifactExistsAsync) is served as the clip, preferring this execution's own rows (loop-back, and
retry, which re-queues the same execution) and then the most recent earlier execution's (a re-run).
- Reused takes reserve and submit nothing. They are removed from the request handed to
IVideoGenerationBudgetGate.ReserveAsync, write no ledger row, count against no cap, and show asreused: true,costUsd: 0,costBasis: releasedin the manifest;meta.clipsReusedandmeta.reusedSavedUsdsummarise them. A fully-reused step never calls the gate. Takes that miss (different prompt, no success, object deleted, lookup error) buy as before, through the same gate. - Default is reuse. Two things force fresh generation:
- The per-execution flag
regenerateVideoClips(defaultfalse) onWorkflowExecutionRequested, set from theexecuteandretryrequest bodies ({ "regenerateVideoClips": true }). It is held in memory (ExecutionRunFlags) for the life of the execution, so a hard process kill degrades to reuse, never to an unrequested spend. The assistant execute path never sets it. - The step config's
loopBackBehavior(Reusedefault |Regenerate):Regeneratepays for fresh takes on every ReviewLoop iteration after the first.
- The per-execution flag
- UI. Run and Retry on a workflow that has a
VideoGeneratestep open a modal offering "reuse previous clips" (default) or "Regenerate (costs money)", showing the sum of the video steps' Max Spend caps as an upper bound (not a quote). The step form exposesloopBackBehavior.
Storage layout and normalization
Each delivered take is uploaded as a bare artifact (no project_files row) under:
projects/{projectId}/agentFiles/generated/{executionId}/step-{stepOrder}-v{n}.mp4
n is the take number on the Inline path (clip-n). On the plan path it is the ordinal of the
clip within the step's dispatched list — 1 for the first clip the step actually bought — which is
not necessarily the clip's own v{n} id. The manifest is what maps a clip id to its object. The
same key is written to the row's storage_key.
Before upload, GeneratedClipNormalizer validates the download with ffprobe (it must contain a
video stream) and re-encodes it to the same canonical settings VideoCompile uses: libx264,
CRF 20, veryfast, yuv420p, constant frame rate at the probed rate (30 fps if unknown), and
AAC audio if the clip has audio. The probe uses VideoEditing:AnalyzeTimeoutSeconds and the
encode VideoEditing:CompileTimeoutSeconds. A failure here ends the take DELIVERY_FAILED.
Configuration
WorkflowEngine appsettings.json, section VideoGeneration (VideoGenerationOptions, bound by
AddVideoGeneration). Override with environment variables such as
VideoGeneration__DailyBudgetUsd.
| Key | Default | Description |
|---|---|---|
VideoGeneration:DailyBudgetUsd | 20.0 | Per-project daily cap in USD. ≤ 0 denies every reservation. |
VideoGeneration:GlobalDailyBudgetUsd | 50.0 | Daily cap across all projects in USD. ≤ 0 denies every reservation. |
VideoGeneration:MaxConcurrentPerProvider | 2 | Takes holding a slot per provider at once; values below 1 act as 1. |
WorkflowHardening:MaxStepRetries (default 3) sets how many attempts a failed VideoGenerate
step gets.
Workflow templates
WorkflowTemplateCatalog ships one template for this feature, video-derush-edit-broll
(AutoCreateOnProject: false, so it is opt-in). It derives a source video, decides the edit, plans
generated b-roll, buys it, and places it in the same encode:
| # | Step | Agent | Notes |
|---|---|---|---|
| 1 | VideoAnalyze | VideoTransform | Source: ProjectFile. The bounded view this emits offers the s{n}/g{n}/t{n} ids. |
| 2 | Agent | VideoStoryEditor | Decides which spans to keep. |
| 3 | Agent | ShotDirector | AgentInputContextMode.FullWorkflow, so it sees step 1's offered ids and step 2's decision. |
| 4 | VideoGenerate | VideoTransform | planSource: "StepRef". |
| 5 | VideoCompile | VideoTransform | enableGeneratedClips: true. |
| 6 | ReviewLoop | VideoReviewAgent | Loops back to step 2, MinScore: 8, MaxIterations: 3. |
The generate step's config:
{
"planSource": "StepRef",
"plan": { "from": "Step", "stepOrder": 3 },
"analysis": { "from": "Step", "stepOrder": 1 },
"model": "MiniMax-H3",
"resolution": "768P",
"durationSeconds": 6,
"ratio": "16:9",
"takes": 1,
"maxClips": 4,
"maxSpendUsd": 5.0,
"pollIntervalSeconds": 10,
"timeoutSeconds": 1800
}
The VideoStoryEditor (step 2) and ShotDirector (step 3) steps both opt out with
CacheMode: Never, since re-asking them is the point of a re-run. The generate step is left at the
default, so an identical re-run does not pay twice.
Every step reference in the template is an explicit StepOrder, never Previous. Two of them
would resolve to the wrong step if they were not:
- On the compile step,
Previouswould resolve to step 4 — theVideoGeneratestep's own output — rather than step 2, the story editor's decision. - On the generate step,
Previouswould resolve to step 3.planwould survive that by accident, since step 3's output is the plan;analysiswould not, because the keyframe ids a planned shot names have to resolve against step 1'sVideoAnalyzeartifact.
The compile step names decision: {from: "Step", stepOrder: 2}, analysisStepOrder: 1,
generatedClips: {from: "Step", stepOrder: 4} and maxGeneratedClips: 4.
Out of scope for phases 3, 4 and 5
Issue #93 §7 lays out the full phasing. Not built yet:
- Higgsfield, and any provider other than MiniMax (wave B / phase 3).
MiniMaxis still the onlyVideoGenerationprovider kind. - Keyframe and image transport (wave B / phase 3). A plan may name a
firstFrame/lastFrame, and the egress gate governs consent, but no frame is sent in this wave — see Egress. - The
SeamBridgeandExtendplacements (wave B / phase 3). Both are accepted vocabulary and degrade explicitly. ScreenContentinto tracked screen inserts (wave C / phase 4).- Dailies, and take selection over the
v{n}t{k}ids (wave C / phase 4). The ids exist so a take-selection pass has something to choose between. - Self-hosted ComfyUI, and the free-form-graph refusal (wave D / phase 5).
- A plan of more than four shots.
MaxClipsis still bounded to 1..4, so a larger plan is trimmed — the excess clips are reportedskippedwithMAX_CLIPS_EXCEEDED— rather than fully generated. The whole plan's estimate is still what the step cap is checked against. - The
Inlinepath is otherwise unchanged. Apart from three things it now shares with the plan path, a phase 1 step behaves exactly as this document originally described: the output projection, the per-clip manifest written toArtifactStorageKey, andmeta.clipsSkipped— all three of them new for anInlinestep in this wave. - MiniMax
callback_url. The step polls; it registers no callback.