Skip to main content

Video Generation

Generate new video clips from a text prompt inside a workflow. You register a MiniMax provider row, add a StepType.VideoGenerate step with a prompt or a plan plus a spend cap, and the step submits one clip per take, polls it, downloads it, normalizes it with ffmpeg, and stores it in the project's object storage. Every submission passes a pre-submit budget gate first.

Issue #93 phase 1 built the provider layer: one provider (MiniMax), one operator-written prompt per step. Phase 2 added the planner. A ShotDirector agent reads the same bounded VideoAnalyze view the story editor decides from and plans the shots; the same deterministic step then buys one clip per planned shot under the same spend caps. A VideoGenerate step therefore has two plan sources: Inline, phase 1's single prompt, and StepRef, a planner step's plan. For the surrounding engine (step types, executors, provider resolution) see CLAUDE.md; for the editing steps that consume a generated clip, including where the compile step places one, see video-editing.md.


Table of Contents​


How a generated clip reaches the rest of a workflow​

VideoGenerateStepExecutor sets the step result's OutputStorageKey to the storage key of the first take that was delivered. A later VideoAnalyze step reads it through its VideoSourceRef like any other rendered output:

  • Source: PreviousStepOutput picks the latest step in the execution with a non-null OutputStorageKey.
  • Source: StepOutput with StepOrder set to the VideoGenerate step picks it explicitly.

See Picking the source video. Clips after the first — further takes on the Inline path, further planned clips on the StepRef path — are stored and listed in the step's clips output, but no VideoSourceRef kind addresses them. On the plan path the manifest is what maps a clip id to its object (see How a planned clip is addressed).

VideoGenerate is a deterministic, non-LLM step: it has no agent of its own and reuses the seeded AgentType.VideoTransform row to satisfy WorkflowStep.AgentDefinitionId, like VideoAnalyze/VideoCompile. It never throws; every failure is a structured JSON output.

Planning the shots​

AgentType.ShotDirector is an LLM agent that plans generated shots from the same bounded VideoAnalyze view the story editor decides from. It is an ordinary Agent step — there is no new step type for it — implemented in inference/src/ReelBolt.WorkflowEngine/Agents/Production/ShotDirectorAgent.cs, with OutputSchemaName = "GeneratedShotPlanOutput". It is seeded as a built-in agent row; its system prompt exists in two places, the class's own DefaultPrompt and DatabaseSeeder.BuiltInAgents, and ShotDirectorPromptConsistencyTests asserts the two stay byte-identical.

It reads two things, which is why a template runs it with AgentInputContextMode.FullWorkflow:

  • the bounded VideoAnalyze view, whose view offers the s{n} (shot), g{n} (silence gap) and t{n} (transcript segment) cut-anchor ids a planned shot may hang off;
  • the story editor's decision, so it can avoid anchoring a shot to a moment the edit cut away.

It emits GeneratedShotPlanOutput — { shots, planRationale }, where each shot is:

{
"purpose": "Cutaway",
"anchor": "s3",
"firstFrame": null,
"lastFrame": null,
"camera": "PushIn",
"duration": "Short",
"prompt": "A slow push across a sunlit desk…",
"reason": "Covers the weakest stretch of picture inside a kept run."
}

ShotPlan is exactly those eight properties and every one of them is a plain string (firstFrame and lastFrame are nullable). purpose is one of Cutaway, SeamBridge, Extend, ColdOpen, EndCard or ScreenContent; duration is one of Short, Medium or Long. An empty shots list is a fully valid outcome — an edit needs no generated shot.

The no-number invariant​

No property anywhere in ShotPlan or GeneratedShotPlanOutput is a numeric or time-bearing type. GeneratedShotPlanOutputInvariantTests is a reflection test that fails if one is added, fails if any ShotPlan property stops being a string, and pins both exact property sets. The planner is therefore structurally incapable of emitting a timestamp, a frame index, a duration in seconds, a resolution, a pixel coordinate or a URL.

The reason matters more than the test. The prompt is the only value the model authors that leaves the building; every other value it emits is either an opaque id that the analysis artifact offered it, or an enum word. First-party C# turns those words into provider parameters: ShotPlanParameterMapper appends an English camera clause to the prompt and snaps the duration word to a number of seconds, while ShotPlanCostEstimator and VideoGenerationPricing do the money. So nothing the model writes is interpreted as a number, a time or a provider parameter.

ShotPlanParameterMapper.CameraWords is the canonical 11-word camera vocabulary. SnapDurationSeconds maps Short to the model's minimum duration, Long to its maximum and Medium to the midpoint, clamped to the model's range in every case. An unknown duration word throws rather than defaulting, because the snapped duration drives the spend estimate and the ledger row; an unknown camera word is simply omitted from the prompt, which is a documented degradation because nothing downstream depends on it.

The prompt sanitizer​

GeneratedPromptSanitizer (ReelBolt.Shared/Inference/GeneratedPromptSanitizer.cs) is the one shared prompt-hygiene pass for model-authored prose; phase 1's private VideoGenerateStepExecutor.SanitizePrompt and its duplicate length constant are gone. DefaultMaxPromptLength is 7000, and MaxPromptLengthFor(InferenceProviderKind) is total over every provider kind (MiniMax 7000, default 7000) so a future kind cannot silently receive a zero cap and disable itself. StripControlCharacters drops every control character except \n, \r and \t.

SanitizedPrompt carries Text, MaxLength, IsEmpty and ExceedsMaxLength, and it does not truncate. An empty or over-length prompt fails closed with INVALID_PROMPT instead, so a provider is never sent a prompt the model did not write. On the Inline path that fails the step; on the StepRef path it degrades that one clip and the step continues.

The planner's tool scope, and the money boundary​

ToolGroupCatalog grants AgentType.ShotDirector exactly ToolGroup.ProjectRead plus ToolGroup.WorkflowControl — the same minimal read-only scope as VideoStoryEditor, MusicSupervisor and SoundDesigner. There is no sandbox tool, no render tool and no project-write tool, and ToolScopingDriftGuardTests.ShotDirector_is_granted_no_sandbox_render_or_spending_tool asserts the grant rather than leaving it to inspection.

The conclusion this is built for: the planner cannot spend money. The planning step is an ordinary agent call, cheap and bounded, and it has nothing that can reach a provider. Only the deterministic VideoGenerate step buys anything, and only within its own config — its MaxSpendUsd and the budget gate below.

MiniMax facts this feature relies on​

Checked against MiniMax's own documentation on 2026-09-28. Each fact links the page that states it.

FactValueSource
Base URLhttps://api.minimax.iocreate
AuthAuthorization: Bearer <api key>create
Create taskPOST /v2/video_generation, body model, content, resolution, duration, ratio; returns task_idcreate
Text-to-video inputcontent holds a single {"type": "text", "text": ...} itemcreate
Prompt lengthat most 7000 characters per text itemcreate
MiniMax-H3768P or 2K; duration integer 4–15 screate
MiniMax-H3-Max480P or 768P (no 2K); duration integer 5–15 screate
Ratio for text-to-videorequired, one of 21:9, 16:9, 4:3, 1:1, 3:4, 9:16; adaptive is not allowedcreate
Prices (pay-as-you-go)MiniMax-H3: 768P $0.08/s, 2K $0.13/s. MiniMax-H3-Max: 480P $0.05/s, 768P $0.08/spricing
Query taskGET /v2/query/video_generation/{task_id} (task_id is a path parameter)query
Task statusesqueued, running, succeeded, failed, cancelledquery
Downloadon succeeded, task.content.url is a time-limited download URL; query again for a fresh onequery
Query windowonly tasks from the last 7 days are queryablequery
Metered usagetask.usage.total_seconds, "input seconds + output seconds", returned only on successquery
Poll cadence10 seconds recommendedguide
Errors400 bad_request, 401 authorized, 402 insufficient_balance, 422 unprocessable_entity, 429 rate_limit, 500 servercreate
Idempotencythe create endpoint documents no idempotency keycreate

The per-model resolutions, durations, ratios and prices above are hard-coded in MiniMaxVideoGenerationClient.Describe (ReelBolt.Shared/Inference/MiniMaxVideoGenerationClient.cs). If MiniMax changes them, that method is the one place to update.

Set up a MiniMax provider​

Admin → Inference Providers → New, or POST /api/v1/inference-providers:

FieldValue
KindMiniMax
CapabilityVideoGeneration
Endpointhttps://api.minimax.io
API keyyour MiniMax API key
Timeoutper-HTTP-request timeout in seconds, applied to submit, poll and download calls
IsDefaultset it to make this the row a step uses when its ProviderId is unset

The step's Model field picks the model; the executor does not read the provider row's model name.

The admin form prefills the endpoint as https://api.minimax.io, the base URL MiniMax documents for the v2 API, when you pick MiniMax.

The API validates the kind/capability pair on create, update and test and returns 400 for:

  • Capability: VideoGeneration with any kind other than MiniMax.
  • Kind: MiniMax with any capability other than VideoGeneration.

There is no env: sentinel for MiniMax. The stored key is decrypted and sent verbatim as the bearer token.

Test a provider without spending​

POST /api/v1/inference-providers/{id}/test (or /test for an unsaved config) never creates a task, so it costs nothing. MiniMaxVideoGenerationClient.ProbeAsync sends one GET /v2/query/video_generation/{random GUID} with the configured key, bounded by a 15-second timeout (the provider row's Timeout, if set lower, cuts it shorter), and classifies the answer:

AnswerResult
401 or 403, or a MiniMax error body whose error.type is authorized_error (any status)ok: false, error "Provider rejected authentication."
500 with error.type server_error and a message starting "record not found" — what MiniMax actually answers for a task id that does not existok: true
A network error, the timeout, 408, 429 or any other 5xxok: false, error "Provider test failed: …" naming what happened
A non-JSON body, or JSON that is neither MiniMax envelope belowok: false, same error form
Any other status whose body is {"task": {...}}ok: true
Any other status whose body is MiniMax's error envelope {"type": "error", "error": {"type": ..., "message": ...}} with an error.type other than authorized_error (for example bad_request_error, "invalid task_id")ok: true

The result's responsePreview (on ok: true and on an auth rejection) carries the status and at most a short, key-scrubbed vendor error.type/error.message, never the raw body. On the saved provider path the outcome is persisted to LastTestAt/LastTestOk/LastTestError.

A passing test means the endpoint answered with MiniMax's query envelope and did not reject the key. MiniMax does not document what the query endpoint returns for a task id that does not exist, but observed behaviour (2026-10-04) is that it checks the key first — an invalid key gets 401 authorized_error — and then answers a valid key's lookup of an unknown id with HTTP 500, server_error, "record not found (1000)". That 500 is why the probe special-cases it: treating every 5xx as a failure made the test fail for every working key. Any other 5xx, including a server_error with a different message, still fails.

Step config reference​

WorkflowStep.VideoGenerateConfigJson (column video_generate_config_json, jsonb) holds a JSON-serialized VideoGenerateStepConfig (ReelBolt.Shared/Workflows/VideoGenerateStepConfig.cs). Property names are camelCase in JSON; enums are strings.

PropertyTypeDefaultNotes
PlanSourceVideoGeneratePlanSourceInlineInline generates Takes clips from InlinePrompt; StepRef generates one clip per planned shot. StepRef is appended after Inline so stored payloads keep their meaning.
InlinePromptstring?nullThe prompt. Required on the Inline path, see below. Ignored on the StepRef path.
PlanExtractInputRef?nullThe planner step's GeneratedShotPlanOutput. Required when PlanSource is StepRef, see below.
AnalysisExtractInputRef?nullThe VideoAnalyze step whose artifact offered the firstFrame/lastFrame ids. Required when any planned shot names one, see below.
Modelstring"MiniMax-H3"MiniMax-H3 or MiniMax-H3-Max. Any other value fails CONFIG_INVALID.
Resolutionstring"768P"Must be one of the model's resolutions, matched exactly (case-sensitive, no trimming).
DurationSecondsint6Seconds per clip on the Inline path. On the StepRef path each shot's own duration word is snapped instead, so this value does not drive the clips — but it is still validated against the model's range, and it is still part of the ledger's request_hash on both paths.
Ratiostring"16:9"Must be one of the model's ratios, matched exactly.
ProviderIdGuid?nullA VideoGeneration provider row. If unset, disabled, deleted or of another capability, the default VideoGeneration row is used.
Takesint1Inline: clips to generate, clip-1..clip-N. StepRef: takes per planned shot, and the clip ids become v{n}t{k}.
MaxClipsint1Author-set hard ceiling on clips for the whole step, validated independently of Takes.
MaxSpendUsddecimal0Step spend cap in USD. No usable default: 0 fails validation.
AllowSourceMediaEgressboolfalseConsent to send frames of the project's own footage to the provider as keyframes. See Egress.
PollIntervalSecondsint10Delay between polls, 1 to 300 inclusive. MiniMax recommends 10.
TimeoutSecondsint1800Per-take polling budget, 60 to 7200 inclusive and at least PollIntervalSeconds, counted from when the take holds a provider slot. A take still unfinished when it elapses ends POLL_TIMEOUT.

Validate() returns every rule that fails; at run time any error ends the step CONFIG_INVALID before the budget gate reserves anything, so nothing is submitted or charged. A missing VideoGenerateConfigJson, invalid JSON or a JSON null also ends the step CONFIG_INVALID.

  • MaxSpendUsd must be greater than 0.
  • Takes must be between 1 and 4 inclusive.
  • MaxClips must be between 1 and 4 inclusive.
  • If MaxClips is in range, Takes must not exceed MaxClips.
  • DurationSeconds must be greater than 0.
  • PollIntervalSeconds must be between 1 and 300 inclusive.
  • TimeoutSeconds must be between 60 and 7200 inclusive.
  • If both are in range, TimeoutSeconds must be at least PollIntervalSeconds; otherwise the take would be paid for and time out before its first poll.
  • Plan is required when PlanSource is StepRef.
  • Plan and Analysis, when set, must be sourced From = Previous or From = Step, and From = Step requires a StepOrder.

MaxClips is a ceiling counted across the whole step, spent in plan order — shot major, take minor. On the StepRef path every clip past it is reported skipped with skipReason: MAX_CLIPS_EXCEEDED and consumes no ledger row, and its bound is still 1..4.

Validate(capabilities) runs after the provider resolves, against Describe(Model), and adds:

  • DurationSeconds must lie within the model's minimum and maximum.
  • Resolution must be a key of the model's price table (ordinal match).
  • Ratio must be one of the model's allowed ratios (ordinal match).

It deliberately adds no StepRef rule: DurationSeconds stays the Inline path's per-clip duration, while the plan path snaps each shot's own duration word, so it is in range by construction.

On the Inline path the executor resolves the provider first — the prompt's length cap is per provider kind — and then sanitizes the prompt:

  • An empty or whitespace-only InlinePrompt fails INVALID_PROMPT.
  • A prompt over the resolved kind's cap (7000 characters for MiniMax) fails INVALID_PROMPT; it is not truncated.
  • Control characters other than \n, \r and \t are stripped first, so a prompt made only of control characters is rejected as empty.

On the StepRef path each planned shot's prompt goes through the same pass, but an empty or over-length prompt fails only that clip.

The Analysis reference is not required by Validate(), because the config alone cannot know what the plan will contain. The executor enforces it: a plan in which any shot names firstFrame or lastFrame and no Analysis is configured ends the step CONFIG_INVALID, naming the offending shots, before anything is reserved.

Save-time validation​

The Inference API rejects a VideoGenerate step whose config would fail Validate() before it is stored, using StepConfigSaveValidator with the executor's own JSON options. A missing or blank videoGenerateConfigJson, invalid JSON and a JSON null are rejected too. Each error reads Step <order> (VideoGenerate): <rule>.

PathRejection
POST and PUT /api/v1/projects/{projectId}/workflows[/{id}]400 {"message": "<errors, one per line>", "errors": [...]}, nothing written
POST /api/v1/projects/{projectId}/workflows/assistant/apply (assistant Apply)same 400 {message, errors}, before the create or update branch writes anything
Assistant ApplyWorkflow tooltool result {"error": "<errors, one per line>", "errors": [...]}, nothing written
Assistant ValidateWorkflow tool and POST .../workflows/assistant/validatea result with valid false and the same errors, so validation never passes a step Apply would reject

Save time checks only the provider-free rules — which now include the structural Plan and Analysis rules, so a plan-sourced config with no Plan is rejected before it is stored. The per-model duration, resolution and ratio checks and the prompt checks need the resolved provider or run in the executor, so they still fail the step at run time.

The web workflow builder (web/lib/utils/video-generate-validation.ts) blocks Save client-side and shows Step N: <first error>. It applies the Validate() rules it can check without a resolved provider, and adds the prompt rules, which the server checks only at run time:

  • InlinePrompt must be non-blank and at most 7000 characters.
  • InlinePrompt must contain no control characters other than \n, \r and \t. The executor strips those at run time instead of failing.
  • A VideoGenerate step with no config is reported as "Max Spend (USD) must be greater than 0", the first error its default config would produce.

Example Inline step config:

{
"planSource": "Inline",
"inlinePrompt": "A slow dolly shot across a sunlit desk with a laptop showing a dashboard",
"model": "MiniMax-H3",
"resolution": "768P",
"durationSeconds": 6,
"ratio": "16:9",
"takes": 2,
"maxClips": 2,
"maxSpendUsd": 1.0
}

That step estimates 2 × 6 s × $0.08/s = $0.96 and passes the $1.00 step cap.

Example plan-sourced step config:

{
"planSource": "StepRef",
"plan": { "from": "Step", "stepOrder": 3 },
"analysis": { "from": "Step", "stepOrder": 1 },
"model": "MiniMax-H3",
"resolution": "768P",
"durationSeconds": 6,
"ratio": "16:9",
"takes": 1,
"maxClips": 4,
"maxSpendUsd": 5.0
}

plan names the ShotDirector step and analysis the VideoAnalyze step whose offered ids resolve any keyframe a shot asks for. durationSeconds is not read for pricing on this path — each shot's own duration word is snapped instead — but it still has to be a duration the model offers, because Validate(capabilities) checks it unconditionally.

The budget gate​

IVideoGenerationBudgetGate.ReserveAsync (WorkflowEngine/Services/VideoGeneration/) checks every cap and writes every take's ledger row as one atomic operation, before anything is submitted.

Pricing is fail-closed. The executor prices each take with VideoGenerationPricing.EstimateClipCostUsd: the model's per-second price for Resolution times DurationSeconds. An unknown model, an unpriced resolution, or a non-positive estimate throws, and the step fails CONFIG_INVALID. An estimate is never 0.

The reservation, in order:

  1. No takes, a take estimate ≤ 0, MaxSpendUsd ≤ 0, or a Guid.Empty step-result id → deny CONFIG_INVALID.
  2. A take whose (ExecutionId, RequestHash) already has a Reserved or Submitted row is returned as existing: not inserted again and not charged again.
  3. Step cap: the sum of every take's estimate (new and existing) > MaxSpendUsd → deny.
  4. Project daily cap: the project's charged spend today + the new takes' estimate > VideoGeneration:DailyBudgetUsd → deny. A cap ≤ 0 denies everything.
  5. Global daily cap: charged spend today across all projects + the new takes' estimate > VideoGeneration:GlobalDailyBudgetUsd → deny. A cap ≤ 0 denies everything. This stops extra projects from multiplying the per-project cap.
  6. Otherwise one Reserved row is inserted per new take and every take is returned in order.

All or nothing. A denial writes no rows and submits nothing; the step fails BUDGET_EXCEEDED (or CONFIG_INVALID) with the gate's reason in meta.error. There is no partial fulfilment.

What counts as spent. "Today" is rows with CreatedAt on or after UTC midnight. Every row except Failed and Cancelled is charged — Reserved, Submitted, Running and Succeeded. A charged row costs max(CostEstimateUsd, ActualCostUsd, 0), so a vendor-reported actual below the estimate never frees headroom, and no row can credit the ledger. Failed and Cancelled rows are released.

Concurrency. The read-then-insert runs under a process-wide lock, inside a transaction, and on PostgreSQL under pg_advisory_xact_lock with a single global key. Two concurrent steps, in the same or different projects and engine processes, cannot both pass a cap only one of them fits under.

Pricing a whole plan​

On the StepRef path the executor prices the whole plan before reserving anything. ShotPlanCostEstimator.EstimateWholePlanUsd sums, over every planned shot, Describe(model).PricePerSecond × snapped duration × takes. Every per-clip price comes from VideoGenerationPricing, the one pricing path, so the estimator cannot drift from the ledger.

It fails closed the same way: an unpriceable resolution or an unresolvable duration word throws and the step ends CONFIG_INVALID rather than contributing a silent 0, because a zero estimate passes every cap. An empty plan is the one legitimate zero — there is genuinely nothing to buy.

The whole-plan estimate is compared against MaxSpendUsd before the gate's reservation. Exceeding it ends the step BUDGET_EXCEEDED with zero submits and zero ledger rows.

A plan larger than MaxClips still has its whole-plan estimate gated. The cap binds the plan the author wrote, not the trimmed subset that would actually be dispatched. The gate is handed only the trimmed take list, so on its own it could never see that the full plan was over budget — which is exactly why this check sits outside it. The gate call that follows (IVideoGenerationBudgetGate.ReserveAsync) then covers the effective, post-trim plan exactly as in phase 1, and meta.planEstimatedUsd is that effective number.

Egress: sending the project's own frames to a provider​

A keyframe is not metadata about the user's footage — it is the user's footage, and once it has been posted to a provider it cannot be recalled. AllowSourceMediaEgress (default false) is the workflow author's explicit consent to send it, and GeneratedShotEgressPolicy decides it per shot.

The executor evaluates that policy against every shot in the plan first. A single refusal fails the whole step with code EGRESS_REFUSED and the policy's reason, with zero submits and zero ledger rows. A partially generated plan is a half-edit the operator cannot use, so the refusal has to be actionable — grant consent, or edit the plan — rather than silently dropping one shot.

With consent granted the step proceeds, but the frame is still not transported. In this wave VideoGenerationRequest has no image field, so a consented keyframe-requiring shot is submitted best-effort from its prompt alone. Each such clip is recorded rather than silently ignored, in the per-clip manifest:

FieldValue
keyframeRequestedtrue
keyframeAppliedfalse
keyframeReasonkeyframe_transport_not_available

meta.keyframesRequested and meta.keyframesNotApplied count them. Transporting the frame is wave B's work — Higgsfield's presigned upload is where it lands — and it is why these fields exist now.

When a shot names a frame at all, the config's Analysis reference is required, and the named id must be one the analysis artifact actually offered the planner. An id that was never offered records keyframeReason: "keyframe_unresolved" on that clip, which is likewise not fatal.

Job lifecycle and crash-resume​

Each take has one row in external_generation_jobs (WorkflowEngine-owned, no foreign keys, so a row outlives its execution, step result and project). Columns: id, execution_id, step_result_id, project_id, clip_id, request_hash, provider_id, remote_job_id, status, cost_estimate_usd, actual_cost_usd, storage_key, created_at, updated_at. status is one of Reserved, Submitted, Running, Succeeded, Failed, Cancelled; the executor never writes Running. step_result_id is the id of the step result for the attempt that reserved the row. A retry that reuses the row (see Attempts and resume) keeps the original attempt's id, so the column does not point at the attempt that last polled. The gate refuses a reservation whose step-result id is Guid.Empty with CONFIG_INVALID.

request_hash is SHA-256 over model, resolution, ratio, duration, take index and the sanitized prompt.

Status per exit path​

Takes run one after another. Every exit writes UpdatedAt and saves with CancellationToken.None, so a stopped step still records its outcome. failureReason in the clip output is the code; detail carries the human explanation.

Exit pathRow status afterwardsfailureReasonCharged?
DeliveredSucceeded, StorageKey set—yes
Vendor reports failedFailedPROVIDER_INSUFFICIENT_BALANCE if the reason contains 402 or insufficient, else PROVIDER_FAILEDreleased
Vendor reports cancelledCancelledPROVIDER_CANCELLEDreleased
Submit returns 4xx other than 408FailedPROVIDER_INSUFFICIENT_BALANCE (402), PROVIDER_AUTH_FAILED (401/403), else PROVIDER_REJECTEDreleased
Submit fails otherwise (5xx, 408, network, malformed 2xx, cancellation) or returns no task idReserved, unchangedSUBMIT_OUTCOME_UNKNOWNyes
Poll throws, or the step is stopped while a submitted take waits or pollsSubmitted, unchangedPOLL_ERRORyes
TimeoutSeconds elapsesSubmitted, unchangedPOLL_TIMEOUTyes
Vendor succeeded with no download URLSubmitted, ActualCostUsd setNO_DOWNLOAD_URLyes
Vendor succeeded, then download, normalize or upload throwsSubmitted, ActualCostUsd setDELIVERY_FAILEDyes
Fresh take never submitted because the step was stoppedCancelledNOT_SUBMITTEDreleased
Take reserved by an earlier attempt with no RemoteJobIdReserved, untouchedRESERVATION_ORPHANEDyes
Any other exceptionCancelled if the take was fresh and never submitted; otherwise unchangedUNEXPECTED_ERRORreleased / yes

Only a vendor-confirmed failure or cancellation, a definite 4xx submit rejection, or a take that never reached the vendor is released. Everything that might have been billed stays charged.

Actual cost​

When the vendor reports succeeded, before downloading, the executor records ActualCostUsd = usage.total_seconds × the model's per-second price for Resolution. A usage of 0 records 0; a missing, null, non-integer or negative usage records null. The executor never copies the estimate into ActualCostUsd.

Attempts and resume​

VideoGenerate is retried like an Agent step: up to WorkflowHardening:MaxStepRetries attempts (default 3), each calling ExecuteAsync again with the same ExecutionId. On the dispatch path the step is Failed only when every dispatched take failed, so that is the only take-driven retry. Every pre-gate refusal is Failed too (CONFIG_INVALID, INVALID_PROMPT, PROVIDER_NOT_FOUND, EGRESS_REFUSED, BUDGET_EXCEEDED, UNEXPECTED_ERROR) and is retried the same way, to the same limit: VideoGenerate is deliberately not one of the deterministic step types capped at a single attempt. Each attempt goes through the budget gate again, within every cap.

On a retry, the gate hands back the earlier attempt's Reserved/Submitted rows as existing:

  • Row has a RemoteJobId: the take is polled only, never re-submitted. This is how a POLL_TIMEOUT or POLL_ERROR take gets another chance, and how a DELIVERY_FAILED take gets a fresh download URL.
  • Row has no RemoteJobId: the earlier attempt may have reached MiniMax without recording the task id. The take ends RESERVATION_ORPHANED with zero submit calls and the row untouched.
  • Failed/Cancelled rows are not returned as existing, so those takes get a new row and a new submission.

Handling an orphaned reservation​

  1. Read the clip's detail: it names the ledger row id and its CreatedAt (UTC).
  2. In the MiniMax console, look for a video task created near that time.
  3. If a task exists, the money was spent and the clip is not in ReelBolt. If none exists, nothing was billed.
  4. Either way the reservation stays charged against the project and global caps until the UTC day rolls over. Do not edit the row to free budget unless you have confirmed no task exists.

Per-provider concurrency​

VideoGeneration:MaxConcurrentPerProvider (default 2, minimum 1) bounds how many takes hold a slot on one provider at once, keyed by provider id. A take acquires the slot before it submits (or, when resuming, before its first poll) and releases it when polling ends, before download and normalization, so a slow download never blocks another submit. The TimeoutSeconds clock starts only once the slot is held. The limiter is per engine process.

MiniMax documents no concurrency limit for the v2 H3 API; the default of 2 is ReelBolt's own conservative choice.

Step output​

A step that reached the gate emits:

{
"clips": [
{ "clipId": "clip-1", "status": "succeeded", "stored": true, "costUsd": 0.48, "costBasis": "actual" },
{ "clipId": "clip-2", "status": "failed", "failureReason": "POLL_TIMEOUT", "detail": "…", "costUsd": 0.48, "costBasis": "estimate" }
],
"meta": {
"totalCostUsd": 0.96,
"actualCostUsd": 0.48,
"unconfirmedCostUsd": 0.48,
"estimatedCostUsd": 0.96,
"clipsGenerated": 1,
"clipsFailed": 1,
"clipsSkipped": 0
}
}

Both plan sources share this projection, so the Inline path's clip-{n} clips and the plan path's v{n} clips are shaped identically. A clip carries {clipId, status, failureReason?, detail?, skipReason?, stored, costUsd, costBasis}. stored is a boolean, not a path — OutputJson reaches later agents under AgentInputContextMode.FullWorkflow, so no storage key may appear in it (see How a planned clip is addressed). shotIndex, take, purpose, anchor and durationSeconds are manifest-only and do not appear here.

costBasis labels each clip's costUsd:

costBasisWhencostUsd
actualActualCostUsd is knownActualCostUsd
estimaterow charged, actual unknownthe row's CostEstimateUsd, never 0
releasedrow Failed or Cancelled0

meta.totalCostUsd sums every clip's costUsd; actualCostUsd and unconfirmedCostUsd sum the actual and estimate clips; estimatedCostUsd sums every take's pre-submit estimate; clipsGenerated, clipsFailed and clipsSkipped count clips by status.

A plan-sourced step adds:

KeyMeaning
planEstimatedUsdThe effective, post-trim plan's estimate — the exact sum of the reservations handed to the gate.
shotsPlannedHow many shots the planner offered, before the MaxClips trim.
keyframesRequestedClips whose planned shot named firstFrame or lastFrame.
keyframesNotAppliedOf those, the ones whose keyframe was not applied.
noteEMPTY_PLAN when the planner legitimately offered no shots at all.

If every take failed, the step status is Failed, meta.failed is ALL_TAKES_FAILED, and ErrorDetails lists clipId:code pairs, e.g. clip-1:POLL_TIMEOUT, clip-2:PROVIDER_FAILED. With at least one delivered take the step is Completed.

A step that fails before the gate reserves anything (CONFIG_INVALID, INVALID_PROMPT, PROVIDER_NOT_FOUND, BUDGET_EXCEEDED, EGRESS_REFUSED, UNEXPECTED_ERROR) emits {"clips": [], "meta": {"failed": "<code>", "error": "<message>"}}.

How a planned clip is addressed​

OutputStorageKey still holds only the first succeeded clip, which is what the UI player and a later VideoAnalyze step read. Every other planned clip is addressed through the per-clip manifest:

projects/{projectId}/agentFiles/video-analysis/{executionId}/step-{order}-generated-clips.json

VideoGenerateStepExecutor writes it and points WorkflowStepResult.ArtifactStorageKey at it. It carries clips[].{clipId, shotIndex, take, purpose, anchor, firstFrame, lastFrame, camera, duration, durationSeconds, status, storageKey, costUsd, costBasis, keyframeRequested, keyframeApplied, keyframeReason, failureReason, detail, skipReason} plus the step's meta, and the compile step resolves generatedClips → that step's ArtifactStorageKey → the manifest.

Clip ids are v{n} in 1-based plan order when Takes is 1, and v{n}t{k} otherwise, so a later take-selection pass has distinct ids to choose between. The Inline path keeps phase 1's clip-{n}.

Why the manifest exists at all. The step's output JSON no longer carries clips[].storageKey. It carries clips[].stored instead, so that no storage path is ever model-visible: the output reaches later agents under AgentInputContextMode.FullWorkflow, and a test asserts the output contains no storageKey and none of projects/, agentFiles/ or outputFiles/, while the manifest behind ArtifactStorageKey carries the real keys.

Because the manifest sits under the video-analysis/ prefix and not outputFiles/, the existing artifact endpoint serves it with no backend change:

GET /api/v1/projects/{projectId}/step-results/{stepResultId}/artifact

It already validates that prefix, so the manifest reads through it exactly like a VideoAnalyze or VideoCompile artifact does.

Where the clips go in the edit​

The compile side gains three appended fields on VideoCompileStepConfig: generatedClips (ExtractInputRef?, default null), enableGeneratedClips (bool, default false) and maxGeneratedClips (int, default 8). Placement derives from each planned shot's own purpose — there is deliberately no per-purpose placement config, because a second config surface would only let the two disagree.

purposeWhat the compile step does
CutawayReplaces picture only over the anchor's window on the OUTPUT timeline. The base audio chain is untouched, so dialogue runs continuously across the cutaway — a J/L-cut.
ColdOpenPrepended as its own concat segment, pairing with ProgramFadeIn.
EndCardAppended as its own concat segment, pairing with ProgramFadeOut.
SeamBridge, Extend, ScreenContentAccepted vocabulary that degrades explicitly to not_supported_in_this_wave rather than being silently misplaced as one of the three above.

A cutaway's window is [anchorStart, min(anchorStart + clipDurationSeconds, anchorEnd)), on the output timeline. Every placement failure degrades per clip and never fails the compile; the one hard config error is enableGeneratedClips with Mode: StreamCopy, which ends the step GENERATED_CLIPS_REQUIRE_REENCODE because a picture replacement and a prepended segment both need a filtergraph.

EnableGeneratedClips: false — the default — is byte-identical to the pre-wave compile path: no extra ffmpeg input, no extra filter, and no new EDL or output key. That backward-compatibility guarantee is load-bearing, the same one the graphics, music, inserts, grade and SFX blocks carry.

For the filtergraph-level detail — how the replacement is gated, how the concat segments are built, and what the EDL's generatedClips node reports — see video-editing.md.

Accepted risks​

  • No vendor idempotency key. MiniMax's create endpoint has none, so exactly-once submission is not achievable client-side. If a submit reaches MiniMax but its task id is never recorded (SUBMIT_OUTCOME_UNKNOWN, or a crash mid-submit), the take is charged at its estimate and never re-submitted, but ReelBolt cannot track or deliver that clip. The operator reconciles through the console (see Handling an orphaned reservation).
  • False orphans err on the charged side. A take reserved but not yet submitted when the process died (for example while waiting for a provider slot) looks the same as one whose submit reached MiniMax. If the same execution runs the step again it is reported RESERVATION_ORPHANED and stays charged, even though nothing was billed.
  • A cache hit validates only the first take. See Step cache.

Step cache​

StepCachePolicy treats VideoGenerate as cacheable. The cache key includes VideoGenerateConfigJson, so the prompt, model, resolution, duration, ratio, takes and provider id (when set) are all part of it. A repeated execution with identical inputs replays the prior Completed result, including a partial one, instead of paying again. A Failed result is never cached.

A plan-sourced step's key also covers the planner's output, transitively: the key is taken over the fully resolved step input, which for a deterministic step type is the preceding step's accumulated output. A plan that changed is therefore a cache miss, and an unchanged one replays without paying twice.

Before serving a hit, the cache HEADs the result's OutputStorageKey, which is only the first delivered take — on the plan path that means only clips[0]'s object. Objects for other clips are not checked, so a hit can reference a deleted clip object.

Setting the step's CacheMode to Never only skips the step cache: identical clips are still reused from the spend ledger (next section). To force fresh, paid takes, run with Regenerate (regenerateVideoClips: true).

Clip reuse on re-run (ledger-level)​

The step cache only hits when the whole resolved input is byte-identical, so a ReviewLoop loop-back (the upstream output changed) or a retry that follows a partial failure would still pay again. Independently of the cache, VideoGenerateStepExecutor therefore reuses clips at the ledger level. Before the budget gate is called, each take's request_hash (model, resolution, ratio, duration, take index, sanitized prompt) is looked up in external_generation_jobs for the same project and provider: a Succeeded row with a StorageKey whose object still exists (HEAD via ArtifactExistsAsync) is served as the clip, preferring this execution's own rows (loop-back, and retry, which re-queues the same execution) and then the most recent earlier execution's (a re-run).

  • Reused takes reserve and submit nothing. They are removed from the request handed to IVideoGenerationBudgetGate.ReserveAsync, write no ledger row, count against no cap, and show as reused: true, costUsd: 0, costBasis: released in the manifest; meta.clipsReused and meta.reusedSavedUsd summarise them. A fully-reused step never calls the gate. Takes that miss (different prompt, no success, object deleted, lookup error) buy as before, through the same gate.
  • Default is reuse. Two things force fresh generation:
    1. The per-execution flag regenerateVideoClips (default false) on WorkflowExecutionRequested, set from the execute and retry request bodies ({ "regenerateVideoClips": true }). It is held in memory (ExecutionRunFlags) for the life of the execution, so a hard process kill degrades to reuse, never to an unrequested spend. The assistant execute path never sets it.
    2. The step config's loopBackBehavior (Reuse default | Regenerate): Regenerate pays for fresh takes on every ReviewLoop iteration after the first.
  • UI. Run and Retry on a workflow that has a VideoGenerate step open a modal offering "reuse previous clips" (default) or "Regenerate (costs money)", showing the sum of the video steps' Max Spend caps as an upper bound (not a quote). The step form exposes loopBackBehavior.

Storage layout and normalization​

Each delivered take is uploaded as a bare artifact (no project_files row) under:

projects/{projectId}/agentFiles/generated/{executionId}/step-{stepOrder}-v{n}.mp4

n is the take number on the Inline path (clip-n). On the plan path it is the ordinal of the clip within the step's dispatched list — 1 for the first clip the step actually bought — which is not necessarily the clip's own v{n} id. The manifest is what maps a clip id to its object. The same key is written to the row's storage_key.

Before upload, GeneratedClipNormalizer validates the download with ffprobe (it must contain a video stream) and re-encodes it to the same canonical settings VideoCompile uses: libx264, CRF 20, veryfast, yuv420p, constant frame rate at the probed rate (30 fps if unknown), and AAC audio if the clip has audio. The probe uses VideoEditing:AnalyzeTimeoutSeconds and the encode VideoEditing:CompileTimeoutSeconds. A failure here ends the take DELIVERY_FAILED.

Configuration​

WorkflowEngine appsettings.json, section VideoGeneration (VideoGenerationOptions, bound by AddVideoGeneration). Override with environment variables such as VideoGeneration__DailyBudgetUsd.

KeyDefaultDescription
VideoGeneration:DailyBudgetUsd20.0Per-project daily cap in USD. ≤ 0 denies every reservation.
VideoGeneration:GlobalDailyBudgetUsd50.0Daily cap across all projects in USD. ≤ 0 denies every reservation.
VideoGeneration:MaxConcurrentPerProvider2Takes holding a slot per provider at once; values below 1 act as 1.

WorkflowHardening:MaxStepRetries (default 3) sets how many attempts a failed VideoGenerate step gets.

Workflow templates​

WorkflowTemplateCatalog ships one template for this feature, video-derush-edit-broll (AutoCreateOnProject: false, so it is opt-in). It derives a source video, decides the edit, plans generated b-roll, buys it, and places it in the same encode:

#StepAgentNotes
1VideoAnalyzeVideoTransformSource: ProjectFile. The bounded view this emits offers the s{n}/g{n}/t{n} ids.
2AgentVideoStoryEditorDecides which spans to keep.
3AgentShotDirectorAgentInputContextMode.FullWorkflow, so it sees step 1's offered ids and step 2's decision.
4VideoGenerateVideoTransformplanSource: "StepRef".
5VideoCompileVideoTransformenableGeneratedClips: true.
6ReviewLoopVideoReviewAgentLoops back to step 2, MinScore: 8, MaxIterations: 3.

The generate step's config:

{
"planSource": "StepRef",
"plan": { "from": "Step", "stepOrder": 3 },
"analysis": { "from": "Step", "stepOrder": 1 },
"model": "MiniMax-H3",
"resolution": "768P",
"durationSeconds": 6,
"ratio": "16:9",
"takes": 1,
"maxClips": 4,
"maxSpendUsd": 5.0,
"pollIntervalSeconds": 10,
"timeoutSeconds": 1800
}

The VideoStoryEditor (step 2) and ShotDirector (step 3) steps both opt out with CacheMode: Never, since re-asking them is the point of a re-run. The generate step is left at the default, so an identical re-run does not pay twice.

Every step reference in the template is an explicit StepOrder, never Previous. Two of them would resolve to the wrong step if they were not:

  • On the compile step, Previous would resolve to step 4 — the VideoGenerate step's own output — rather than step 2, the story editor's decision.
  • On the generate step, Previous would resolve to step 3. plan would survive that by accident, since step 3's output is the plan; analysis would not, because the keyframe ids a planned shot names have to resolve against step 1's VideoAnalyze artifact.

The compile step names decision: {from: "Step", stepOrder: 2}, analysisStepOrder: 1, generatedClips: {from: "Step", stepOrder: 4} and maxGeneratedClips: 4.

Out of scope for phases 3, 4 and 5​

Issue #93 §7 lays out the full phasing. Not built yet:

  • Higgsfield, and any provider other than MiniMax (wave B / phase 3). MiniMax is still the only VideoGeneration provider kind.
  • Keyframe and image transport (wave B / phase 3). A plan may name a firstFrame/lastFrame, and the egress gate governs consent, but no frame is sent in this wave — see Egress.
  • The SeamBridge and Extend placements (wave B / phase 3). Both are accepted vocabulary and degrade explicitly.
  • ScreenContent into tracked screen inserts (wave C / phase 4).
  • Dailies, and take selection over the v{n}t{k} ids (wave C / phase 4). The ids exist so a take-selection pass has something to choose between.
  • Self-hosted ComfyUI, and the free-form-graph refusal (wave D / phase 5).
  • A plan of more than four shots. MaxClips is still bounded to 1..4, so a larger plan is trimmed — the excess clips are reported skipped with MAX_CLIPS_EXCEEDED — rather than fully generated. The whole plan's estimate is still what the step cap is checked against.
  • The Inline path is otherwise unchanged. Apart from three things it now shares with the plan path, a phase 1 step behaves exactly as this document originally described: the output projection, the per-clip manifest written to ArtifactStorageKey, and meta.clipsSkipped — all three of them new for an Inline step in this wave.
  • MiniMax callback_url. The step polls; it registers no callback.