Video Editing
Automatic derushing and editing of real, uploaded or rendered video files — silence and shot
detection, optional ASR transcription, an LLM editorial decision, and a frame-accurate ffmpeg cut,
extended with multi-clip cutting, motion graphics, background music, and a multi-agent "edit room"
deliberation — implemented as three new deterministic-or-orchestrating workflow step types plus
five new built-in LLM agents (VideoStoryEditor, MotionGraphicsPlanner, VideoReviewAgent,
MusicSupervisor, VideoEditDirector) and one deterministic placeholder agent (VideoTransform).
This document is the reference for that feature; for the surrounding workflow engine (step types,
executors, agents in general) see CLAUDE.md.
Table of Contents
- The three-stage shape
- The id-anchored decision contract
- Where artifacts live
- Config reference
- Multiple source clips
- Playing spans in the editor's order
- Using only part of a source
- Keeping part of a long shot
- Output format
- Program audio on silent footage
- Compile warnings
- One registered render per run
- Takes: the latest render first
- Scene/visual analysis (Phase 1)
- Vision captioning (Phase 2)
- Transcription (ASR)
- Transcripts from clips with no speech
- Motion graphics (Phase 3)
- Background music
- Cutting to the beat
- Seam transitions and the program envelope
- Semantic visual dimensions (Phase 4)
- Tracked screen inserts (Phase 5)
- Non-planar surfaces and multi-plate compositing
- Silhouette matting
- Surviving occlusion
- Memory: every branch starts at zero
- A clip's last frame and the next clip's first
- The edit room
- The shared room infrastructure
- The graphics room
- Color grading
- Sound effects
- Generated clips (b-roll)
- Security: why ffmpeg is not in the sandbox
- Explicitly not built
The three-stage shape
┌─────────────────┐ ┌────────────────────────┐ ┌──────────────────┐
│ StepType. │ │ StepType.Agent │ │ StepType. │
│ VideoAnalyze │────▶│ AgentType. │────▶│ VideoCompile │
│ │ │ VideoStoryEditor │ │ │
│ deterministic, │ │ LLM, structured output │ │ deterministic, │
│ ffmpeg + ASR │ │ (VideoEditDecision- │ │ ffmpeg │
│ │ │ Output) │ │ │
└─────────────────┘ └────────────────────────┘ └──────────────────┘
emits {view, meta} emits an ordered list of resolves ids to
+ a full analysis opaque "Keep" id spans frame-accurate
artifact — never a timestamp times, encodes
-
StepType.VideoAnalyze(ReelBolt.Shared/Workflows/VideoAnalyzeStepConfig.cs,WorkflowEngine/Execution/StepExecutors/VideoAnalyzeStepExecutor.cs) — deterministic, non-LLM. Resolves the source video, probes it, detects silence gaps and shot/scene changes with ffmpeg, optionally transcribes the audio, assigns every detected item a short opaque id, and emits:- a bounded
{view, meta}prompt envelope (the same shapeStepType.Extractproduces, so downstreamExtractsteps can compose over it unchanged), persisted tooutput_json - the full, non-truncated analysis document (
VideoAnalysisArtifact), persisted separately as a storage-key artifact
- a bounded
-
StepType.Agent+AgentType.VideoStoryEditor— an ordinary Agent step. No new step type was introduced for the editorial decision:AgentStepExecutoralready provides structured output (ChatResponseFormat.ForJsonSchema<T>()), per-agent provider resolution, retry-with- feedback, and tool scoping, so reusing it is a straight win. The agent is given the bounded view from step 1 and decides which ids to keep. -
StepType.VideoCompile(ReelBolt.Shared/Workflows/VideoCompileStepConfig.cs,WorkflowEngine/Execution/StepExecutors/VideoCompileStepExecutor.cs) — deterministic, non-LLM. Loads the full analysis artifact from step 1, resolves the agent's chosen ids to exact[start, end)times, validates and normalizes the resulting cut list, frame-quantizes it, and encodes the edited video with ffmpeg. Optionally (Phase 3,EnableGraphics), also composites motion-graphics overlays during that same encode — see Motion graphics (Phase 3).
An optional fourth stage — StepType.Agent + AgentType.MotionGraphicsPlanner — can sit between
steps 2 and 3, planning overlays from placement candidates step 1 derived alongside its usual cut
anchors. It follows the exact same shape as step 2 (an ordinary Agent step, no new step type). See
Motion graphics (Phase 3) for the full design.
Both VideoAnalyze and VideoCompile follow the same discipline ExtractStepExecutor
established: they never throw. Every failure mode — an unresolvable source, a missing
transcription provider, an unknown id, overlapping spans — is represented as a structured failure
in the step's JSON output, because WorkflowStepResult.OutputJson is a jsonb column and an
unhandled exception there takes down the whole execution's SaveChangesAsync.
Picking the source video
VideoAnalyzeStepConfig.Source is a VideoSourceRef, one of three kinds:
| Kind | Resolves via | Use case |
|---|---|---|
PreviousStepOutput (default) | The latest completed step result in this execution with a non-null output storage key | "Edit whatever the previous step in this workflow rendered" — e.g. chain straight off a Remotion render step |
StepOutput + StepOrder | That specific step's WorkflowStepResult.OutputStorageKey | Reference an earlier step explicitly, regardless of what runs in between |
ProjectFile + ProjectFileId | project_files.storage_key | Edit a raw video/audio file the user uploaded |
This three-way split exists because Remotion render outputs are not ProjectFile rows —
ReactRemotionSandboxTools.RenderVideoAndUploadToStorage uploads directly to
projects/{projectId}/outputFiles/{executionId}/{name} and records only
WorkflowStepResult.OutputStorageKey, never inserting into project_files. A source model that
could only address ProjectFile ids would be unable to express "edit the video this workflow just
rendered" — the primary use case.
The id-anchored decision contract
The rushcut invariant: the story-editor agent can never emit a timestamp. This is not a prompt
instruction the model could drift away from — it is structural. VideoEditDecisionOutput
(ReelBolt.Shared/Data/OutputSchemas.cs) has no numeric, TimeSpan, or DateTime property
anywhere in it:
public class VideoEditKeepSpan
{
public string FromId { get; set; } = ""; // e.g. "t7" or "s2" — first offered id to keep
public string ToId { get; set; } = ""; // last offered id to keep, inclusive
public string Reason { get; set; } = ""; // prose only
}
public class VideoEditDecisionOutput
{
public List<VideoEditKeepSpan> Keep { get; set; } = new(); // ordered, non-overlapping
public string EditRationale { get; set; } = "";
public string SuggestedTitle { get; set; } = "";
}
There is no "remove" list — everything not covered by a Keep span is cut. A second,
"remove ids" representation would create two ways to express the same edit and a resolution order
to get wrong, for zero benefit.
This is enforced in CI, not just by convention: VideoEditDecisionOutputInvariantTests
(inference/tests/ReelBolt.WorkflowEngine.Tests/VideoEditDecisionOutputInvariantTests.cs) uses
reflection to assert that neither VideoEditDecisionOutput nor VideoEditKeepSpan exposes any
int/long/float/double/decimal/TimeSpan/DateTime/DateTimeOffset property (nullable
variants included). If a future change adds e.g. a StartSec field "to make things easier", this
test fails the build.
Ids, and why "offered" is a stricter check than "exists"
Every shot, silence gap, transcript segment, and word in the full analysis artifact gets a
deterministic id, assigned by index: s{n} shots, g{n} silence gaps, t{n} transcript
segments, w{n} words (words are persisted for audit/phase-2 subtitle export but never offered to
the model — segments are the offered granularity in v1). The artifact also records
OfferedIds: exactly which ids were actually included in the bounded view shown to the model,
which can be a strict subset of every id in the full artifact once VideoAnalyzeStepConfig's
view-budget trimming drops trailing items.
VideoCompileStepExecutor rejects any FromId/ToId that is not in OfferedIds — not merely
present somewhere in the full artifact. This closes a subtle hole: without it, a model could
reference an id it was never actually shown (e.g. one dropped by trimming), which would resolve to
a real time in the artifact but not one the model ever saw or reasoned about.
How compile resolves ids to frame-accurate times
VideoCompileStepExecutor, in order — never trusting a model-emitted number, because there is
none in the schema to trust:
- Load the full analysis artifact via the referenced
VideoAnalyzestep'sArtifactStorageKey(AnalysisStepOrder, orAnalysisStepResultIdto resolve a prior execution's artifact — scoped to the same project only, see below). - Reject any id not in
OfferedIds(UNKNOWN_ID). - Map each id to
[StartSec, EndSec)from the full artifact. - Normalize: sort by start; require strictly increasing, non-overlapping spans (reordering kept
spans is out of scope for v1 — a violation fails the step with a precise diagnostic that
becomes retry feedback); coalesce adjacent spans; apply
PrePaddingMs/PostPaddingMs; clamp to[0, durationSec]; drop spans shorter thanMinSegmentMs; enforceMaxSegments. - Frame-quantize on the exact rational fps read from ffprobe's
r_frame_rate(e.g.30000/1001), using integer arithmetic throughout — never a roundeddouble— so accuracy never depends on the source's frame rate being a whole number. - Evaluate
Expect(includingMinRetainedRatio, which refuses an edit that discards nearly everything the source contained). - Write the EDL artifact (the resolved cut list, for audit), then encode.
AnalysisStepResultId lets a VideoCompile step reference a VideoAnalyze result from a
different execution — the first-class mechanism behind the approval-by-composition workflow
below. It is resolved only against step results whose WorkflowExecution.ProjectId matches the
current execution's project; a cross-project reference is refused.
The cut's rationale is joined to its OWN Keep span, never by loop index
TimelineItem.AgentRationale on each cut-N item is the Reason of the VideoEditKeepSpan that
cut actually descends from. That join used to be by loop index i —
decision.Keep[i].Reason, else the first non-empty reason anywhere in Keep — which is only
valid while the raw Keep list survives the span pipeline 1:1, and step 4 above makes sure it does
not: spans are coalesced, spans below MinSegmentMs are dropped, and the list is capped
at MaxSegments. After any of those, index i names a different Keep entry than the one the cut
came from, so a cut routinely showed another cut's reason, and a cut with no reason of its own
silently inherited the first reason in the list.
Instead, every span now carries ResolvedSpan.OriginKeepFromIds: the FromId(s) of each
Keep[i] it descends from, threaded through the raw-span tuple pipeline from the moment the span is
first resolved (new[] { span.FromId }) to the frame-quantized ResolvedSpan — including the
no-valid-fps early return, so a span keeps its origins even when quantization cannot run.
CoalesceAdjacent concatenates both sides' origin ids in order (previous side first) rather than
keeping one, so a coalesced span still names every Keep entry it covers. BuildVideoTrack builds a
keepByFromId dictionary once, outside the loop, and resolves the rationale by looking each origin id
up in it.
A span whose origin list is null or empty — and a span whose origin ids all resolve to a blank reason
— falls back to decision.EditRationale, the edit-level rationale. The old index lookup and its
"first non-empty reason anywhere" fallback are gone: there is no path left that can attribute one
cut's reason to another. See The timeline manifest for the item fields this
rationale lands in.
Approval: composition, not a suspend/resume gate
There is no blocking mid-execution approval gate in v1 — WorkflowExecutorService.ExecuteAsync
has no AwaitingApproval status or persisted suspend point to resume from; building one is a
whole initiative of its own. Instead:
- Run a workflow containing only
VideoAnalyze+ theVideoStoryEditoragent step. Inspect the proposed edit in the execution UI (the "Edit decision list" panel, or theVideoAnalyzestep's own bounded view/artifact). - Run a second workflow or execution containing only
VideoCompile, pointed at the first run's analysis artifact viaAnalysisStepResultId.
The compiled video needs no new playback UI: writing OutputStorageKey under the outputFiles
prefix makes it appear in GET /outputs and in the existing execution-page <video> player
automatically, exactly like a Remotion render.
Where artifacts live
| Artifact | Storage key | Referenced by |
|---|---|---|
Full analysis JSON (shots + silences + words + segments + OfferedIds + provenance — can be several MB) | projects/{projectId}/agentFiles/video-analysis/{executionId}/step-{order}-analysis.json | WorkflowStepResult.ArtifactStorageKey on the VideoAnalyze result |
Bounded prompt view {view, meta} (≤ MaxOutputChars, default 24,000) | WorkflowStepResult.OutputJson (jsonb) | Agent input / step history, same as any other step |
| Resolved EDL (the compile step's own cut list, for audit) | projects/{projectId}/agentFiles/video-analysis/{executionId}/step-{order}-edl.json | WorkflowStepResult.ArtifactStorageKey on the VideoCompile result |
| Timeline manifest (the finished edit as named studio tracks + per-item AI provenance) | projects/{projectId}/agentFiles/video-analysis/{executionId}/step-{order}-timeline.json | WorkflowStepResult.TimelineStorageKey on the VideoCompile result (see The timeline manifest) |
| Edited mp4 | projects/{projectId}/outputFiles/{executionId}/{OutputFileName} | WorkflowStepResult.OutputStorageKey and a new project_files row (category outputFiles, mime video/mp4) |
| Working files (intermediate segment files, caption keyframes) | {VideoEditing:ScratchPath}/{executionId}/{stepId}/ (default /var/tmp/reelbolt-video) | Nothing — deleted in a finally block once the step completes |
| Extraction result, for resume | projects/{projectId}/agentFiles/video-analysis/{executionId}/step-{order}-extract.v1.json | Nothing — the key is derived from (project, execution, step order) by VideoAnalysisExtractArtifact.BuildStorageKey |
| The extracted full-track WAV | projects/{projectId}/agentFiles/video-analysis/{executionId}/step-{order}-src{n}.wav | VideoAnalysisExtractSource.AudioStorageKey, carried inside the extract document |
ArtifactStorageKey is a separate column from OutputStorageKey on purpose.
OutputsController.ListOutputs treats any step result whose OutputStorageKey starts with
projects/{id}/outputFiles as a playable video, and the execution UI renders a <video src=…>
for it. Putting the analysis JSON or EDL there would make the UI try to play JSON as a video.
Keeping them in a separate column under a separate agentFiles/video-analysis prefix avoids that
entirely, and lets a dedicated endpoint validate against a narrower, non-playable prefix:
GET /api/v1/projects/{projectId}/step-results/{stepResultId}/artifact
Mirrors OutputsController.DownloadOutput's guards exactly (project exists → caller owns it →
step result resolved through WorkflowExecution.ProjectId) plus one more: the storage key must
start with projects/{projectId}/agentFiles/video-analysis — never outputFiles — so this
endpoint can never be coerced into serving a playable render or another project's object.
Registering the edited video as a ProjectFile (unlike Remotion render outputs, which are
not) is deliberate: it makes an edited video re-editable via VideoSourceKind.ProjectFile and
visible in the project's file list. It's created with SummaryStatus = Done and
IndexingStatus = NotIndexed so the text summarizer/vector chunker never touches a binary — see
the R22 fix below.
The EDL is written and uploaded exactly once per compile. VideoCompileStepExecutor builds
edl.json, writes it to scratch and uploads it under its final key before it starts encoding; the
uploaded document deliberately does not carry timelineStorageKey inline. An earlier version
re-wrote that same local file and re-uploaded the same storage key after the manifest existed,
purely to stamp that one field onto it — doubling the EDL's upload cost on every compile, for a
value every consumer already had elsewhere, and making the "one EDL artifact per run" invariant
untrue. The key reaches callers through the step's OutputJson (timelineStorageKey) and
WorkflowStepResult.TimelineStorageKey only; nothing reads it from inside the EDL. The single
write+upload is unconditional and runs before the encode, so the ENCODE_FAILED path still returns
its edlStorageKey — a failed encode keeps the cut list describing what it tried to cut.
The persisted extraction (resume)
A VideoAnalyze step runs in two phases: a deterministic extraction phase — stage, probe,
silence, shots, the full-track WAV, audio levels, frame grids and everything read off them, chroma
tracking, look grouping, candidate listing, beat facts, the caption keyframes — and an inference
phase (ASR, vision captioning, subject location, object tracking, the derush pre-pass, input
screening, the artifact and the bounded {view, meta} envelope) that is the only thing which calls
a provider.
The extraction phase's result is persisted as a versioned document
(analysis-extract.v1.json, VideoAnalysisExtractArtifact) at the key in the table above, and a
later attempt of the same step loads it before doing anything else — one HEAD, then a GET.
A hit skips extraction entirely, so an ASR outage, a vision provider timing out on the last of ten
captions, or any other inference-side failure costs the inference phase and nothing else. That is
what a remote target needs in order not to re-dispatch a runner's extraction because the LLM half
failed.
What a resume recovers, and what it deliberately does not. The analysed media is re-staged from
the object store key the document records (for a still image the rendered clip, for an
author-trimmed source the trimmed window — the media every id and time in the document is relative
to), and the extracted WAV from its published key. Both are downloads, never decodes. The caption
keyframes are not published (publishing them would duplicate the user-visible PersistKeyframes
output and collide with the -keyframes/ path that output owns), so a resumed attempt whose session
is gone re-solves the keyframes it cannot find — one seek plus one frame each, bounded by
MaxCaptionedShots — and records any that still fail exactly where extraction would have. A session
that survived the failed attempt (a remote/runner session) finds them present and pays nothing.
The WAV publish is load-bearing for the resume. It stays non-fatal, because in place inference reads the session file and a transient upload error must not fail a step that would otherwise succeed — but a missing key now has a defined consequence instead of a silent one: a document whose source has an extracted WAV but no published key is refused, and extraction runs again. That is the honest answer, because the alternative is a resume that reads no audio and reports an empty transcript for a video that has speech.
The document never becomes a source of truth. It is addressed by (project, execution, step
order), so "it is still there" is not the same question as "it is still valid": a loader re-resolves
every source reference and refuses the document when an upstream output has changed (a
ReviewLoop loop-back), when the schema version is not the one this build writes, or when the
source count no longer matches. Every refusal degrades to extracting again, which is exactly the
behaviour the step had before the document existed. Losing the document costs a re-extraction and
never correctness, so a failure to persist it is logged and never fails the step.
The timeline manifest
Every successful VideoCompile also emits a timeline manifest: a TimelineManifest
(ReelBolt.Shared/Workflows/TimelineManifest.cs) describing the finished edit as named studio
tracks (video, inserts, graphics, dialogue, sfx, music), each holding TimelineItems
positioned on the output timeline. Its key is published on
WorkflowStepResult.TimelineStorageKey — a third column, separate from OutputStorageKey
and ArtifactStorageKey for exactly the reason above (the manifest is JSON, and the execution UI
plays anything under outputFiles). It is indexed on workflow_step_results, so
TimelinesController.ListTimelines is a single indexed predicate rather than a scan over
OutputJson, and the value is mirrored onto workflow_step_cache_entries and restored onto the
reconstructed result on a cache hit — a cached VideoCompile still lists its timeline.
This document covers the v1 manifest — a read-only description of a finished compile — and the
pipeline that produces it; the v2 editable document (EditTimeline) a client opens, edits and
saves back, together with its op catalogue, sync protocol and trust boundary, is
docs/editor.md, and the manifest is what a v2 timeline is seeded from.
TimelineItem.AgentName/ProviderName/StepResultId are reported facts, never inferred. They
name the agent that really authored the item's decision, the inference provider that agent's step
resolved to, and that step's newest WorkflowStepResult id — each read from the execution's own
step graph and database rows. TimelineProvenance/TimelineProvenanceContext
(WorkflowEngine/Services/Video/TimelineProvenance.cs) splits the work: the executor resolves
the facts (it owns the DbContext scope and the step graph), and TimelineManifestBuilder — which
stays pure, synchronous and I/O-free — stamps them. The builder takes a single trailing, defaulted
provenance parameter, so every pre-existing call site is unchanged.
Each track stamps its items from its own slot of that context, keyed off the same
VideoCompileStepConfig reference the decision itself was resolved from:
Track (TimelineTrack.Type) | Provenance slot | Config reference |
|---|---|---|
video (cuts) | Decision | Decision |
graphics | Graphics | GraphicsPlan |
music | Music | MusicPlan |
sfx | Sfx | SfxPlan |
video (cuts) — fallback only | ColorGrade | ColorGradePlan |
inserts, dialogue | (no slot — never stamped) | deterministic: no agent authored them |
Cuts win over the colour grade. A video item takes the Decision provenance whenever it
resolves; the ColorGrade slot fills in only when the decision provenance is unavailable and
the item really carries a ColorGrade value. A colour grader did not author an ungraded cut, so
stamping grade provenance onto one would be precisely the fabricated label this provenance exists to
remove. The two are never merged: one item names exactly one author.
A room step resolves its director, not its placeholder. EditRoom/GraphicsRoom/
ColorGradeRoom steps point their own AgentDefinitionId at the non-LLM VideoTransform
placeholder that merely satisfies the non-nullable FK, so stamping it would have labelled every
room-authored cut VideoTransform. The real author comes from the room's own config surface, exactly
as RoomStepExecutorBase resolves it: the configured DirectorAgentDefinitionId when present, else
the room kind's built-in director type (VideoEditDirector, MotionGraphicsDirector,
ColorGradeDirector). A room config that is present but undeserializable returns neither — whether
an override was configured is then unknowable, so the track degrades to null rather than guessing the
built-in director.
All three fields null is a real answer, not a defect. Provenance resolution is best-effort by
contract and can never fail a compile: an unresolvable step reference, a step absent from the
execution, a missing AgentDefinition row, or a service scope with no IInferenceProviderResolver
all leave that track's items unstamped — exactly what the manifest looked like before provenance
existed. The three facts are independent, so a partial answer is normal: a step result not yet
persisted loses only StepResultId, and a provider-resolution failure loses only ProviderName,
keeping the agent name. Null therefore means deterministic (inserts, dialogue) or unresolved
(everything else), never "guessing failed so here is a plausible label". TimelineInspector
(web/components/editor/TimelineInspector.tsx) reads it the same way — provenance is reported by
the backend, never re-derived on the client from which other fields happen to be populated, so an
item with no agentName simply shows no author rather than an invented one.
Uploaded video/audio never reaches the text pipeline (R22)
ProjectFilesController.Upload (and the folder-move and bulk-reindex endpoints) skip both
summarization-queue enqueue and vector-indexing enqueue when the file's MIME type starts with
video/ or audio/, setting SummaryStatus = Done and IndexingStatus = NotIndexed immediately
instead of the normal Pending → summarizer/chunker pipeline. Without this, a large uploaded mp4
or wav would be handed to ProjectFileIndexingConsumer, which reads the object as UTF-8 text
before chunking it — wasted compute at best, a very large in-memory string at worst.
Config reference
VideoAnalyzeStepConfig
| Field | Default | Notes |
|---|---|---|
Version | — | Config schema version |
Source | — | VideoSourceRef — see Picking the source video. Ignored when Sources is non-empty |
Sources | null | Multi-source addition — IReadOnlyList<VideoSourceRef>. When non-empty, the AUTHORITATIVE list of source clips analyzed into ONE merged artifact; null/empty (default) falls back to treating [Source] as a one-element list — see Multiple source clips |
DetectSilence | true | ffmpeg silencedetect |
SilenceThresholdDb | -34.0 | |
MinSilenceMs | 350 | |
DetectShots | true | ffmpeg scene-change detection |
SceneThreshold | 0.30 | |
MaxShotSec | 0 | Split every detected shot longer than this into sub-shots carrying part — see Keeping part of a long shot. 0 splits nothing (byte-identical) |
Transcription | Optional | Off / Optional / Required — see Transcription |
TranscriptionProviderId | null | Explicit override; otherwise resolved via the default Transcription-capability provider |
Language | null | Passed through to the ASR provider |
WordTimestamps | true | |
MaxAsrChunkBytes | 20,000,000 | ASR chunk size cap (~10.4 min of 16kHz mono s16 WAV); long audio is split at silence-boundary-aligned chunks, never mid-word |
MaxDurationSeconds | 1800 | Guardrail, checked before any decode |
MaxInputBytes | 2,000,000,000 | Guardrail, checked before any decode |
MaxOutputChars | 24,000 | Prompt-view budget — trimming drops whole trailing items and re-serializes, never truncates mid-JSON (same rule as ExtractStepConfig) |
MaxViewSegments | 400 | |
MaxSegmentTextChars | 160 | |
AnalyzeVisuals | true | Phase 1: one low-res grid ffmpeg pass + pure C# analyzer — see Scene/visual analysis (Phase 1) |
VisualSampleFps | 2.0 | Grid sample rate, clamped downward by MaxVisualSampleFrames for long videos |
VisualGridWidth / VisualGridHeight | 32 / 18 | Downscaled grid resolution the analyzer runs against |
MaxVisualSampleFrames | 4000 | Caps the grid buffer size: effectiveFps = min(VisualSampleFps, MaxVisualSampleFrames / durationSec) |
StillMotionThreshold | 0.02 | Per-frame motion (0..1) below which a moment counts as "still" |
MinStillWindowMs | 400 | Minimum duration for a still run to be reported as a StillWindow |
MaxStillWindowsPerShot | 3 | Longest still windows kept per shot |
DetectLetterbox | true | D6: implemented (Phase 4), free from the grid data Phase 1 already samples — see Semantic visual dimensions (Phase 4). Default changed from false; may under-report soft/gradient letterbox edges |
DetectSharpness | false | D-adjacent "sharpness": implemented (Phase 4) but costs one extra native-resolution ffmpeg invocation per measured shot (capped by MaxSharpnessShots), so it stays opt-in unlike the other free dimensions — see Semantic visual dimensions (Phase 4) |
AnalyzeAudioLevels | true | Phase 1: WavRmsSampler over the WAV already extracted for transcription, or extracted fresh if transcription is off |
DetectNearDuplicates | true | Phase 1: near-duplicate/best-take grouping via FrameGridAnalyzer.GroupDuplicates |
DuplicateSimilarityThreshold | 0.90 | Minimum signature similarity (0..1) for two shots to be grouped |
DuplicateWindowShots | 20 | Single-linkage grouping only compares a shot against the previous N shots (multi-take shots are temporally adjacent) |
VisualDetail | Compact | None / Compact / Full — how much per-shot visual/audio detail the bounded view includes; degrades toward None before any item is ever dropped — see below |
MaxViewDuplicateGroups | 20 | Caps view.duplicateGroups |
AnalyzeColorGrading | true | Phase 4, D1-D3 (colour temperature, tone curve, saturation character) — free, three adds and one array increment inside a pixel loop that already runs — see Semantic visual dimensions (Phase 4) |
DetectLookGroups | true | Phase 4, D4 look grouping (view.lookGroups, ids k{n}) — free, O(shots²) over six floats, cheaper than the 576-float duplicate grouping already running |
LookSimilarityThreshold | 0.88 | Minimum look similarity (0..1) for two shots to share a look group; lower than DuplicateSimilarityThreshold since look distance is a weighted six-vector, not a 576-float grid comparison |
MaxViewLookGroups | 12 | Caps view.lookGroups |
Vision | Off | Phase 2: Off / Optional / Required — vision-LLM shot captioning, off by default (unlike Transcription) — see Vision captioning (Phase 2) |
VisionProviderId | null | Explicit override; otherwise resolved via the default Vision-capability provider |
CaptionSelection | PerDuplicateGroup | PerDuplicateGroup / LongestShots / EvenlySpaced — which shots get captioned; round-robins across source clips when more than one is analyzed (Phase 4) |
MaxCaptionedShots | 50 | Hard cap on vision chat-completion calls across the whole step (genuinely step-wide since the Phase 4 vision hoist — see Vision captioning (Phase 2)), not a target: every shot under MinCaptionShotSeconds is excluded before this cap even applies. Default raised from 24 |
MinCaptionShotSeconds | 1.0 | Shots shorter than this are never selected for captioning |
KeyframeMaxWidth | 512 | Max width (px) of the extracted keyframe JPEG (or contact sheet — see KeyframesPerShot) sent to the vision model; never upscaled |
VisionTimeoutSeconds | 120 | Aggregate wall-clock budget for the whole captioning pass (not per-shot) |
MaxCaptionChars | 320 | Caption summary field is truncated to this length |
KeyframesPerShot | 1 | Phase 4: frames combined into ONE contact-sheet keyframe per captioned shot, clamped 1..3. 1 (default) is a single mid-shot still, byte-identical to the pre-Phase-4 vision path — see Semantic visual dimensions (Phase 4) |
PersistKeyframes | false | Phase 4: now wired — when true, each captioned shot's keyframe JPEG is uploaded to storage under the video-analysis/{executionId}/step-{n}-keyframes/ prefix (a persist failure never costs the caption itself). When false (default), keyframes stay scratch-only and are deleted with the rest of scratch space |
EmitOverlayPlacements | false | Phase 3: derives deterministic overlay-placement candidates (view.placements) from each shot's Phase 1 region data — see Motion graphics (Phase 3) |
MaxPlacementsPerShot | 2 | Top-N regions (by Suitability) offered per shot |
MaxPlacements | 40 | Hard cap on placements across the whole artifact; lowest-suitability candidates dropped first |
MaxTimeSlicesPerRegion | 3 | When a region's chosen time window is long enough to hold more than one distinct overlay moment, split it into up to this many non-overlapping, evenly-spaced sub-windows instead of offering every overlay the same window — see Motion graphics (Phase 3) |
MaxSharpnessShots | 24 | Phase 4: step-wide ceiling on sharpness measurements when DetectSharpness is on — genuinely step-wide like MaxCaptionedShots, not per source. Costs one extra ffmpeg invocation per measured shot |
AnalyzeMusicBeats / MusicBeatTrackProjectFileId / MaxBeatAnalyzedTracks | false / null / 6 | Measure tempo, beat, bar, drop and outro of offered music tracks and of the track the edit is cut to, for the editor (view.musicTracks[].beat, view.music) and for the compile's BeatSync — see Cutting to the beat |
OfferMusicTracks | false | Enumerates every audio/* project file as an m{n} music-track candidate (view.musicTracks) for a downstream AgentType.MusicSupervisor step — project-level, not per-source. See Background music |
MaxMusicTracks | 20 | Caps view.musicTracks |
OfferSfxClips | false | Enumerates every audio/* project file as an x{n} SFX-clip candidate (view.sfxClips) for a downstream AgentType.SoundDesigner step — project-level, sharing one ListFilesAsync call with OfferMusicTracks when both are on. See Sound effects |
MaxSfxClips | 40 | Caps view.sfxClips |
DetectInsertRegions | false | Phase 5: chroma-plate quad tracking for tracked screen inserts — one extra medium-res grid ffmpeg pass per source + pure C# (ChromaQuadTracker); tracks offered as r{n} ids (view.insertRegions) — see Tracked screen inserts (Phase 5) |
MaxInsertPlatesPerFrame | 1 | Chroma plates tracked SIMULTANEOUSLY per frame — the multi-plate capability. 1 is the original largest-component-only behaviour exactly; clamped 1..16 — see Non-planar surfaces and multi-plate compositing |
MeasureInsertCurvature | true | Measure each plate's silhouette (refined corners + edge curvature). The refined corners improve the FLAT composite too; false reproduces the pre-existing composite exactly — see Non-planar surfaces and multi-plate compositing |
SolveInsertChromaKey | true | Grid-search each tracked plate's own colorkey parameters so the compile step can clip the insert to the plate's real per-frame silhouette — rounded bezel corners, a camera notch, a hand crossing the screen — see Silhouette matting |
InsertRegionColor | "green" | "green"/"blue"/"magenta" — matched entirely in C# channel-ratio space, NEVER an ffmpeg value; unknown values fall back to green |
InsertSampleFps | 0 (auto) | Sample rate of the dedicated tracking pass. Auto = the source's own frame rate, or the smallest exact divisor of it ≤ 30 fps that fits the frame cap and memory budget; a positive value is an explicit override (clamped 0.5..30). See Sub-pixel corners at the source's frame rate |
InsertGridWidth / InsertGridHeight | 0 / 0 (auto) | Tracking grid resolution (clamped 64..640 / 36..360). 0 derives it from the SOURCE's probed aspect (~320px long axis), so grid-space and source-space distances agree — see "The tracking grid preserves the source's aspect ratio". Set BOTH to override |
MaxInsertSampleFrames | 3000 | Clamps effective tracking fps downward for long videos, exactly like MaxVisualSampleFrames |
MinInsertRegionAreaRatio | 0.004 | Minimum fraction of frame area a chroma component must cover to count as a plate |
MinInsertRegionSeconds | 1.0 | Tracks shorter than this are dropped |
MaxInsertRegions | 8 | Cap on offered tracks across the whole artifact (longest kept) |
Expect | null | Optional structural checks (MinShots, MinTranscriptSegments, MaxSilenceRatio, MinShotsWithVisuals) |
VideoCompileStepConfig
| Field | Default | Notes |
|---|---|---|
Version | — | Config schema version |
Decision | — | ExtractInputRef (reused verbatim from StepType.Extract) — only From = Previous or From = Step are valid here |
AnalysisStepOrder | — | Which VideoAnalyze step's artifact to resolve ids against |
AnalysisStepResultId | null | Cross-execution override — resolve a prior run's artifact instead of this execution's; scoped to the same project |
Mode | Reencode | Reencode (frame-accurate select/aselect filtergraph) or StreamCopy (fast, lossless, but cuts snap to keyframes) |
PrePaddingMs / PostPaddingMs | 80 / 120 | |
MinSegmentMs | 250 | Spans shorter than this are dropped |
MaxSegments | 200 | Above ~64 segments the filtergraph is written to a scratch file and passed via -filter_complex_script to avoid argv length limits |
AllowKeyframeSnapping | false | Must be true to use Mode = StreamCopy |
OutputFileName | "edited.mp4" | Sanitized to a safe character set with a forced extension. When left at the default and RegisterProjectFile is on, the project file is registered as "{workflow name} - run {n}.mp4" (plus (revision {k}) for a review-loop re-compile) instead — see Compile warnings for the naming rule. A custom name is kept verbatim |
OutputFormat | Source | Source (canvas = the largest kept source), Vertical (1080x1920), Square (1080x1080), Landscape (1920x1080). A fixed format requires Mode = Reencode. See Output format |
OutputFit | Crop | How a clip whose aspect differs from a fixed format is fitted: Crop (scale to cover, centre crop) or Pad (scale to fit, filled per CanvasFill). Ignored for Source |
TargetDurationSec | 0 | The length the customer asked for; 0 = none. A longer cut is trimmed to it (Target length); reported as length in the step output, with a target_length_missed warning when off by more than 20% |
VideoCodec | "libx264" | Allowlisted (libx264, libx265, libvpx-vp9) — config is workflow-author-supplied, not model output, but still reaches ffmpeg argv. Whatever the codec, a re-encoded output is always 4:2:0 (yuv420p), standard range, with the MP4 index up front (+faststart), and H.265 is tagged hvc1: without an explicit pixel format the encoder followed the filtergraph to full-range or even 4:4:4 output (H.264 "High 4:4:4 Predictive" once tracked inserts composited in planar RGB), which played on a desktop and was refused by phones and WhatsApp |
AudioCodec | "aac" | Allowlisted (aac, libmp3lame, copy) |
Crf | 20 | Clamped 0..51 |
Preset | "veryfast" | Allowlisted (ultrafast … veryslow) |
RegisterProjectFile | true | Registers the compiled video as a re-editable ProjectFile row |
GraphicsPlan | null | Phase 3: ExtractInputRef (Previous/Step only) — which step's resolved MotionGraphicsPlanOutput to apply. null = no graphics looked up. See Motion graphics (Phase 3) |
EnableGraphics | false | Phase 3: applies the resolved graphics plan during the same encode. false (default) is byte-identical to the pre-Phase-3 compile path. Requires Mode = Reencode |
MaxOverlays | 20 | Cap on applied overlays; excess dropped (recorded in the graphics block) |
OverlayShortMs / OverlayMediumMs / OverlayHoldMs | 1500 / 3000 / 6000 | Milliseconds an overlay stays on screen, keyed by the model's Duration word (Short/Medium/Hold) |
OverlayFadeMs | 300 | Fade-in/fade-out duration at each end of an overlay's on-screen window |
OverlayFontSizePct | 5 | Percent of frame height; clamped 2..12 at execution time |
OverlayBoxHeightPct | 16 | Percent of FRAME height the drawn overlay box (drawbox/drawtext background, or the box a rendered-asset overlay is stretch-scaled into) occupies — a compact accent strip, not the named safe-zone band's own height. Clamped 6..40, never exceeding the band's own height |
OverlayBoxWidthPct | 82 | Percent of the named band's own WIDTH the drawn overlay box occupies, centered. Clamped 30..100 |
OverlayFontColor | "white" | Allowlisted (white/black/yellow/#RRGGBB) — reaches ffmpeg's filter string, so validated like VideoCodec |
OverlayBoxColor | "[email protected]" | Allowlisted ([email protected], [email protected], [email protected], none) |
MaxOverlayTextChars / MaxOverlaySubtextChars | 80 / 60 | Sanitized-text truncation budget (OverlayTextSanitizer) |
MusicPlan | null | Background music: ExtractInputRef (Previous/Step only) — which step's resolved MusicPlanOutput to apply. null = no plan looked up (deterministic MusicTrackProjectFileId path, or no music, is used instead). See Background music |
MusicTrackProjectFileId | null | A specific audio/* project file to use as the music track — the deterministic path (no agent required), and also the fallback when MusicPlan is unresolvable/invalid or names an unoffered track id |
BeatSync | Off | Beat / Bar: snap every cut to the music's beat or bar by trimming span tails, and cut on the drop — see Cutting to the beat. Off is byte-identical |
EnableMusic | false | Applies the resolved music track during the same encode. false (default) is byte-identical to the pre-music compile path. Requires Mode = Reencode and AudioCodec != "copy" |
MusicDucking | SpeechEnvelope | Off / SpeechEnvelope — SpeechEnvelope lifts the music during non-speech windows via a deterministic keyframed volume envelope; Off is a constant ducked bed throughout |
MusicFitPolicy | LoopToFit | LoopToFit (loops via -stream_loop -1 to fill the whole edit, then trims to its exact length) / PlayOnce (plays once, trimmed to its own length if shorter than the edit) |
MusicFadeInMs / MusicFadeOutMs | 1500 / 2500 | Fade duration at the start/end of the music track's own play window |
MusicBedQuietDb / MusicBedBalancedDb / MusicBedFeatureDb | -26 / -20 / -14 | Bed level (dBFS), keyed by the model's Intensity word (Quiet/Balanced/Feature). Clamped [-40, -6]. When the output has no dialogue and no narration (silent footage, or MuteSourceAudio) the music is the whole soundtrack and plays at -1 dB regardless, through a −2 dBFS peak limiter (alimiter, lookahead compensated) because a mastered track already decodes over full scale (music.bedBasis: "sole_soundtrack") |
MuteSourceAudio | false | Drop the clips' own sound from the program: only music, narration and sound effects are heard. For generated clips that carry their own (often unwanted) audio. audio.reason reads source_audio_muted when it removed real audio |
MusicDuckLightDb / MusicDuckNormalDb / MusicDuckHeavyDb | -6 / -11 / -18 | Attenuation (dB, below the bed) applied while dialogue is present, keyed by the model's Ducking word. Clamped [-30, 0] |
MusicDuckRampMs | 400 | Linear gain ramp (ms) INSIDE each lift window — the music is never above the ducked level exactly at a window boundary |
MinMusicLiftWindowMs | 1200 | Non-speech windows shorter than this (and shorter than twice the ramp) are never lifted at all |
MaxMusicLiftWindows | 12 | Caps the volume-envelope expression's length; excess windows dropped, longest first, then re-sorted chronologically |
MusicLiftMergeMs | 400 | Lift windows closer together than this are merged into one |
TransitionPolicy | Off* | VideoTransitionPolicy — gates the deterministic seam-transition system; Off is byte-identical to the pre-transition compile path. The video-derush-edit* templates opt in with "Auto". See Seam transitions and the program envelope |
AudioSeamRampMs | — | Duration of the audio-only declick ramp treatment at a cut seam |
SoftCutMs | — | Duration of a soft-cut (brief cross-blend, shorter than a full dissolve) treatment |
DissolveMs | — | Duration of a full crossfade-dissolve treatment |
DipToBlackMs | — | Duration of a dip-to-black treatment (fades to black, then to the next shot) |
DipCutMs | — | Duration of a dip-cut treatment (a very brief dip, shorter than a full dip-to-black) |
MaxTransitionMs | — | Hard cap on any single transition's duration, regardless of treatment |
MaxTransitionRatioPct | — | Caps a transition's duration as a percentage of the SHORTER of its two neighbouring segments, so a transition can never eat a meaningful fraction of either one |
MaxTransitionSegments | — | Above this many segments in the compile, transitions are skipped entirely (every seam falls back to a hard cut) — a filtergraph-buffering guardrail, not a quality knob; see Seam transitions and the program envelope |
SectionBreakGapMs | — | Minimum silence-gap duration at a seam for the rule table to treat it as a section break rather than an ordinary mid-sentence cut |
ProgramFadeInMs / ProgramFadeOutMs | — | Video fade-in/fade-out duration at the very start/end of the whole compiled program (distinct from any inter-cut transition) |
ProgramAudioFadeInMs / ProgramAudioFadeOutMs | — | Audio fade-in/fade-out duration at the very start/end of the whole compiled program, tracked independently of the video program fade |
EnableInserts | false | Phase 5: composites the plan's chosen screen inserts (corner-pinned via ffmpeg's per-frame-animated perspective filter) during the same encode. false (default) is byte-identical to the pre-inserts compile path. Requires Mode = Reencode (INSERTS_REQUIRE_REENCODE); inserts are read from the SAME GraphicsPlan-referenced MotionGraphicsPlanOutput (EnableGraphics itself need not be on) — see Tracked screen inserts (Phase 5) |
InsertSurface | Auto | Whether a plate may be composited with a piecewise-projective MESH warp instead of one corner pin. Auto decides per plate from measured curvature; Planar never meshes; Mesh always tries — see Non-planar surfaces and multi-plate compositing |
MinInsertCurvature | 0.01 | Measured curvature a plate must exceed under Auto to earn a mesh, as a fraction of its own size |
InsertMeshTolerancePx | 0.6 | Target worst-case chord error per mesh cell — this is what sets the cell count |
MaxInsertMeshCells | 24 | Ceiling on rows * cols for one insert's mesh |
EnableInsertMatte | true | Clip each composited insert to the plate's own per-frame silhouette, so anything passing in FRONT of the plate occludes the insert instead of being painted over. A plate with no solved key composites with the plain quad mask regardless — see Silhouette matting |
InsertReflections | true | Screen-blend the plate's own glare and reflection streaks over a matted insert, so the new screen content carries the reflections the filmed glass had. A clean plate contributes nothing — see Reflections |
MinInsertConfidence | 0.5 | Tracks below this confidence are dropped (confidence_below_threshold) rather than composited badly |
MaxInserts | 3 | Cap on applied inserts; excess dropped in plan order |
MaxInsertExprKeyframes | 1500 | Per-insert cap on corner keyframes baked into the perspective expressions (uniform downsample; clamped 2..5000). Past 96 keyframes the series is written as a log-depth balanced tree, which is what lets a track sampled at the source's frame rate reach the encode undecimated |
InsertOverscan | 0.02 | Fractional outward expansion of the tracked quad about its centroid, hiding the plate's edge fringe under the insert (clamped 0..0.1) |
SfxPlan | null | Sound effects: ExtractInputRef (Previous/Step only) — which step's resolved SfxPlanOutput to apply. null = no SFX looked up. Unlike music, deliberately NO deterministic no-agent fallback field — see Sound effects |
EnableSfx | false | Mixes the resolved SFX cues into the output audio during the same encode. false (default) is byte-identical to the pre-SFX compile path. Requires Mode = Reencode and AudioCodec != "copy" (SFX_REQUIRES_REENCODE/SFX_REQUIRES_AUDIO_REENCODE) |
MaxSfxCues | 8 | Cap on applied cues; excess dropped in plan order (max_cues_exceeded) |
MaxSfxCueSeconds | 4.0 | Hard cap on any single cue's play window — a long file misused as a cue is trimmed, never a de-facto bed |
SfxSubtleDb / SfxNormalDb / SfxStrongDb | -18 / -12 / -6 | Cue gain (dBFS), keyed by the model's Volume word. Clamped [-40, 0] |
SfxLeadMs / SfxLagMs | 150 / 150 | How far a Timing: "Lead"/"Lag" cue fires before/after its anchor's output-timeline start moment |
SfxFadeOutMs | 120 | Declick fade-out at the end of each cue's play window, capped at half the window |
ColorGradePlan | null | Color grading: ExtractInputRef (Previous/Step only) — which step's resolved ColorGradePlanOutput to apply (a solo Agent(Colorist) step or a ColorGradeRoom step; both emit the same shape). null = no grade looked up. See Color grading |
EnableColorGrade | false | Applies the resolved colour grade (a first-party eq/colorbalance/colorlevels/hue chain keyed by the plan's enum words — ColorGradeFilterBuilder) during the same encode, before overlays/inserts. false (default) is byte-identical to the pre-grade compile path. Requires Mode = Reencode (COLOR_GRADE_REQUIRES_REENCODE); every plan-level failure degrades to "no grade applied" |
Expect | null | Optional structural checks (MinOutputSeconds, MaxOutputSeconds, MinRetainedRatio default 0.15, MaxRetainedRatio) |
KeepOrder | SourceOrder | AsListed plays every kept span in the editor's listed order, even within one clip — see Playing spans in the editor's order. SourceOrder is byte-identical to before |
* TransitionPolicy and the twelve fields above it (AudioSeamRampMs through
ProgramAudioFadeOutMs) belong to a sibling in-flight change adding the deterministic
seam-transition system to VideoCompileStepExecutor/VideoCompileStepConfig; as of this doc edit
they are not yet present on the shipped VideoCompileStepConfig type in this worktree, only
referenced by name in the video-derush-edit* templates' seeded config JSON and in this section, so
their defaults are intentionally left blank above pending that merge.
The video-derush-edit template
An opt-in workflow template (AutoCreateOnProject: false, seeded in
ReelBolt.Shared/Workflows/WorkflowTemplateCatalog.cs) demonstrating the full pipeline:
VideoAnalyze (Source: PreviousStepOutput) → Agent(VideoStoryEditor) → VideoCompile
(Decision: Previous, AnalysisStepOrder: 1). Its literal seeded JSON is deserialization-tested
against the real config types in
WorkflowTemplateCatalogConfigDeserializationTests.cs, so a future field-name drift between the
template and the config records it targets fails CI rather than a live workflow run.
The video-derush-edit-graphics template
A fourth opt-in template (AutoCreateOnProject: false), extending video-derush-edit with Phase
3 motion graphics end to end: VideoAnalyze (Source: ProjectFile, EmitOverlayPlacements: true)
→ Agent(VideoStoryEditor) → Agent(MotionGraphicsPlanner) → VideoCompile
(Decision: Step 2, AnalysisStepOrder: 1, EnableGraphics: true, GraphicsPlan: Step 3).
Decision/GraphicsPlan reference their source steps explicitly by StepOrder rather than
Previous, since Previous relative to the compile step would resolve to the
MotionGraphicsPlanner step's output, not the story editor's decision. Deserialization-tested the
same way as video-derush-edit.
The video-derush-edit-music template
A fifth opt-in template (AutoCreateOnProject: false), extending video-derush-edit with
background music instead of graphics: VideoAnalyze (Source: ProjectFile,
OfferMusicTracks: true) → Agent(VideoStoryEditor) → Agent(MusicSupervisor) → VideoCompile
(Decision: Step 2, AnalysisStepOrder: 1, EnableMusic: true, MusicPlan: Step 3). Same
explicit-StepOrder rationale as video-derush-edit-graphics above (Previous relative to the
compile step would resolve to the MusicSupervisor step's own output, not the story editor's
decision). See Background music. Deserialization-tested the same way as the
other two templates.
Multiple source clips
VideoAnalyzeStepConfig.Sources (plural — IReadOnlyList<VideoSourceRef>) lets one VideoAnalyze
step analyze several source clips (e.g. multiple takes, camera angles, or B-roll of the same scene)
into ONE merged artifact, which a single VideoStoryEditor decision and a single VideoCompile
step can then cut across. It is a strict superset of the original single-clip behavior: Sources
null/empty (the default) is treated as a one-element [Source] list, so every existing
persisted/template config — which only ever set the singular Source — keeps deserializing and
behaving byte-identically. A one-element Sources list behaves identically to the equivalent
single-Source config too; there is no separate "N=1" code path anywhere in this addition.
Analysis: independent per-clip passes, merged into one global id space
VideoAnalyzeStepExecutor analyzes each clip independently via the exact same deterministic
per-source pipeline it always ran (silence/shot detection, transcription, Phase 1/2/3/4 analysis),
processed sequentially, never in parallel, one clip's local file at a time. The pre-decode
guardrails (MaxDurationSeconds, MaxInputBytes) are enforced per source clip, not summed
across all of them.
Each clip's own ids restart at s0/g0/t0/w0/p0/d0. VideoAnalyzeStepExecutor.OffsetId
then remaps every local id into ONE globally-unique id space via a running per-id-kind offset
(assigned contiguously across every source file, in analysis order), and every shot/silence
gap/segment/word/placement/duplicate-group is tagged with the SourceIndex of the clip it came
from. The bounded view surfaces this as a "src" index on every offered item (e.g. "src": 0), so
the VideoStoryEditorAgent prompt can tell the model which clip each id belongs to and instruct it
to freely alternate between clips across successive Keep spans — picking whichever clip has the
best material for each moment is the whole point of offering more than one. The one hard rule: a
single Keep span's FromId and ToId must both come from the SAME clip, since a span is a
contiguous run within one physical file, never a bridge across two files — cross-clip edits are
expressed as a SEQUENCE of single-clip Keep spans instead.
Per-source provenance (VideoAnalysisProvenance) is aggregated into one artifact-level record via
AggregateProvenance: applied/degraded flags become true if ANY source applied/degraded that
stage — an artifact-wide OR, not a per-source breakdown. Full per-source provenance detail is
deliberately out of scope for this addition; a single source passes through unchanged (Count == 1
short-circuits).
VideoAnalysisArtifact.Sources records one VideoAnalysisSourceInfo per analyzed clip, in
source-index order — the storage key VideoCompileStepExecutor must download to physically cut
from that clip, and the per-clip VideoAnalysisMedia (duration/fps/dimensions) every clip-aware
computation (frame quantization, padding clamps, graphics geometry) must use instead of a single
artifact-wide Media. null only for a true legacy artifact produced before this field existed —
the one case VideoCompileStepExecutor still re-derives the source storage key the old way, by
walking the VideoAnalyze step's own config. Every artifact produced by the current executor
populates Sources with at least one entry, even for a single source, so the top-level Media and
Sources[0].Media always agree for that case.
Compile: which ids can pair, and how the cut is resolved
VideoCompileStepExecutor resolves each Keep span's FromId/ToId to [Start, End) times plus
the SourceIndex recorded against that id — never trusted from the model, since there is no source
field on VideoEditKeepSpan for it to get wrong in the first place:
- Spans across clips — a single span whose
FromIdandToIdresolve to differentSourceIndexvalues is SPLIT, when its ids run forward in the offered order, into one span per clip (SplitAcrossSources: consecutive same-source runs, the first keeping the span's ownFromIdso its transition still applies), reported asspanNormalization.splitAcrossSources. Every still image is its own source, so "the last beat of one photo to the first of the next" is an ordinary thing for an editor to write, and failing the workflow on it was worse than the obvious reading. A span whose ids run BACKWARDS across clips still fails withMIXED_SOURCE_SPAN, before any normalization runs. - Ordering/overlap/coalescing only ever compares a span against the immediately PRECEDING span in list order; when that neighbor belongs to a DIFFERENT source clip, there is no shared timeline to be "out of order" or "overlapping" on, so the check (and coalescing) is simply skipped at that boundary. A single-source config's spans are always same-source neighbors, so this reduces to exactly the original single-timeline behavior.
- Padding/clamping clamps each span against ITS OWN source clip's duration (
GetSourceMedia), never a single artifact-wide duration. - Frame-quantization uses each span's OWN source clip's exact rational fps, never a single artifact-wide fps.
- Retained ratio (
Expect.MinRetainedRatio) sums only the DISTINCT clips actually referenced by the resolved cut list, not every clip the step merely analyzed — for a single source this sum has exactly one term, so it is byte-identical to before this addition. - Only the DISTINCT source clips actually referenced by the resolved cut list are downloaded —
never every clip the
VideoAnalyzestep analyzed. - The "canonical" clip every multi-source encode normalizes toward (scale/pad/fps for video, sample rate/channel layout for audio) is the LARGEST kept source by pixel area, ties going to the clip kept first — still deterministic, independent of clip count or offered-id ordering. It used to be the first kept span's clip, which shrank a narrated edit's 1920x1080 b-roll to the 1344x768 of the generated clip that happened to open it. Every geometry computation (overlays, inserts, censor masks) reads the canvas through the same per-clip fit the encode applies; see Output format.
MULTI_SOURCE_REQUIRES_REENCODE — when the resolved cut list references more than one distinct
source clip, Mode = StreamCopy is refused as a hard, pre-encode config error (the same discipline
GRAPHICS_REQUIRE_REENCODE already established for a different Reencode-only combination):
losslessly concatenating independently-encoded files has no correctness-preserving stream-copy
equivalent, since ffmpeg's concat filter/demuxer both require matching codec parameters across
inputs that separately-encoded source files are not guaranteed to share, and normalizing them first
is itself a re-encode.
Why multi-source encoding is a separate method
EncodeReencodeMultiSourceAsync is a deliberately separate method from the original
EncodeReencodeAsync — which stays completely untouched, and is still used for every single-source
compile, so a single-source (or single-clip-in-practice) compile's ffmpeg argv/behavior stays
byte-identical to before this addition. It is not a generalization of the single-source method
because ffmpeg's select filter always emits one input's own matched ranges in THAT INPUT'S OWN
chronological order — it cannot express a Keep-span order that jumps between clips arbitrarily.
Only the concat filter, fed one small pre-trimmed clip PER SPAN in the exact order they should
play, can. Each span becomes its own trim/atrim branch off the correct ffmpeg input index for
that span's own source clip, normalized to the canonical frame size/rate/audio format concat
requires, then concatenated in Keep order.
Audio-less source clips
Real B-roll/stock footage routinely ships with no audio stream at all. VideoCompileStepExecutor
probes every distinct referenced source once (cheap) before encoding and threads the result through
as a per-source sourceHasAudioByIndex map:
- If NOT ONE referenced clip has an audio stream and nothing is mixed in, the whole compiled output drops audio entirely — there is nothing to preserve, so an all-silent track would add nothing. When narration or sound effects ARE mixed in, a silent base is synthesized instead; see Program audio on silent footage.
- Otherwise (a mix of audio-having and audio-less clips),
concat's own stream-count contract (every concatenated segment must carry the samea=count) is satisfied per span: a span whose own source clip has audio gets its realatrimbranch; a span whose source clip has no audio instead gets a synthesized, matching-duration silence branch (ffmpeg'sanullsrcsource filter, already natively 48kHz/stereo, so it needs no extra-iinput oraformat). Every span with real dialogue keeps its real dialogue — only the audio-less span(s) carry synthesized silence. The step's output JSON records which resolved-span indices got synthesized silence (audio.syntheticSilenceSegmentCount/syntheticSilenceSegmentIndices), so this is never a silent surprise the way an unreported drop would be. - Whether the FINAL OUTPUT has any dialogue audio at all (
hasDialogueAudioInOutput—isMultiSource ? anySourceHasAudio : sourceHasAudio) also feeds Background music's ducking decision: there is no point planning silence-gap ducking windows against dialogue that will not exist in the output.
Playing spans in the editor's order
By default a compile plays the spans of one source clip in that clip's own order, whatever order the
editor listed them in: within each run of consecutive same-clip spans they are sorted by start time
and overlaps merged (outputSummary.spanNormalization reports what moved). Spans of different clips
already played in the listed order. So "open on the strongest shot" was impossible when that shot was
filmed last — story order was source order.
VideoCompileStepConfig.KeepOrder (VideoKeepOrder, append-only, persisted by name) changes that:
| Value | Behaviour |
|---|---|
SourceOrder (default) | Exactly the behaviour above — byte-identical |
AsListed | Every span plays in the editor's listed order, within one clip too |
Under AsListed nothing is sorted. Two rules keep the cut well-formed:
- No moment plays twice.
RemoveEarlierCoverageremoves from each span every part an EARLIER span (in list order) of the same clip already covers. A span split in two by an earlier one keeps both pieces (in source order); one covered entirely disappears. It runs again after padding, so a padded boundary never replays a few frames of an earlier span. - Only forward continuations merge.
CoalesceInListOrdermerges a span into the one before it only when it starts exactly where that one ends in the same clip. A span that jumps back is a deliberate reorder and stays its own span.
select/aselect (the single-clip encode) can only play a file forwards, so when any span of a
clip plays before an earlier part of the same clip (CountOutOfSourceOrder > 0) the compile takes
the segmented trim + concat encode, even for one source. With Mode = StreamCopy that is a hard
config error, REORDER_REQUIRES_REENCODE. An AsListed decision that happens to be in source order
produces exactly the default cut. The seam planner treats a backward jump within one clip like a cut
between clips (no "removed gap" is measured for it).
The output summary carries, only under AsListed:
"keepOrder": { "mode": "AsListed", "outOfSourceOrderSpans": 1, "overlapTrimmedSpans": 0 }
The decision schema is unchanged — VideoEditDecisionOutput was always an ordered list, and its
no-timestamp invariant tests are untouched. The VideoStoryEditor prompt (both copies, the agent's
DefaultPrompt and DatabaseSeeder) gained a separate section, "Playing spans out of source order
(only when the brief asks)": list spans in the order they should be seen when the brief asks for a
different order, never list a moment twice, and keep source order otherwise. On a compile left at
SourceOrder a reordered list is simply put back in source order, so the prompt can never break a
cut.
Still images
An image is a source like any clip: a ProjectFile (or step output) whose key ends in .png, .jpg,
.jpeg, .webp, .bmp, .tif/.tiff or .avif — or, with no extension, whose probe reports an
image codec and no duration — is turned into video once, in VideoAnalyzeStepExecutor.RenderStillAsync,
and the clip is recorded as the source exactly like an author-trimmed window, so the editor, captions,
music, grading and the compile all work on it unchanged.
The clip (StillImageClip) is the photo at its own aspect (long side at most 1920, even sides),
upscaled 2x before zoompan so the moves do not judder, with three MOVES back to back — a push in
(8%), a pull back out, and a pan along the long side at 1.12x — each two 1.5 s beats, 9 s in all, no
audio. Each beat is declared as an author scene (CompositionSceneManifest, IsComposition: false),
so the analysis offers six shots per photo, named push in (1/2) … pan right (2/2) (pan down for
a tall photo). The FIRST beat's caption leads with what the photo shows — ONE vision caption of the
image when Vision is on and a provider resolves, "A still photo." otherwise, capped at 200
characters — and every other beat says "Same photo as this source's first shot" and its move: six
copies of one caption per photo pushed a 13-source view over its budget, and the view's only remedy
(detail down to None) stripped every caption, move names included. The editor therefore chooses the
motion by which shots it keeps and its length by keeping one beat (1.5 s) or both (3 s) — all by id.
The VideoStoryEditor prompt (both copies) has a "Photos" section: the six shots are
alternatives, so keep ONE move per photo (a first run kept all six beats of every photo, 74 s for a
25 s brief). meta.stillImages lists the still sources. InSec/OutSec on an image are ignored; a render failure
fails the step with STILL_IMAGE_FAILED.
Using only part of a source
A workflow author can use only a window of a source clip: VideoSourceRef gained InSec and
OutSec (seconds of the original file; either may be null for "from the start"/"to the end";
appended to the positional record, so every stored payload deserializes unchanged and both null is
byte-identical). The web clip list shows Use from / Use until (seconds) on every clip.
VideoAnalyzeStepExecutor.TrimSourceAsync handles it once per source, right after the download:
- The window is validated against the file's probed length (
InSec >= 0, at least 0.1 s long, the start inside the file; anOutSecpast the end is clamped to it). Anything else fails the step withSOURCE_TRIM_INVALIDbefore any decode. - The window is cut out frame-accurately (input seek + re-encode,
BuildTrimArgs:libx264 -crf 16, the source's own frame rate, its audio when it has any) and uploaded as the step's own artifact (video-analysis/{executionId}/step-{n}-src{i}-window.mp4). - Everything else — probe, silence and shot detection, transcription, the visual grid, insert and
object tracking — runs on the WINDOW, and the artifact records the window as the source
(
VideoAnalysisSourceInfo.StorageKey/Media), plusTrimInSec/TrimOutSecfor the record.
So every id, time and track is relative to the window's own start, and no later step needs to know a
trim happened: VideoCompileStepExecutor downloads the recorded window and maps ids through it
exactly as for any source. Author-declared scenes (a sidecar next to a render) describe the whole
piece, so a trimmed source detects its shots instead. The envelope's meta.sourceWindows
([{ src, inSec, outSec }]) lists the windows, present only when a source was trimmed. Trimming
needs the video tool runner (always registered in the engine; a construction without it fails with
SOURCE_TRIM_UNAVAILABLE), and SOURCE_TRIM_FAILED reports an ffmpeg failure.
Keeping part of a long shot
An editor that may only name ids could keep a long shot whole or drop it whole — a great shot with a
bad second at its end had to go. VideoAnalyzeStepConfig.MaxShotSec (default 0, off) splits every
detected shot longer than that (values under 1 s act as 1 s) into n = ceil(length / MaxShotSec)
sub-shots (ShotSplitter.Split). Each split starts at the even division and moves, within 30% of a
part either way, to:
- the stillest moment there — the mean absolute luma change between consecutive frames of the
Phase 1 visual grid (
ShotSplitter.MotionCurve, smoothed over three samples). The grid is sampled once, before the split, and reused by the visual analysis, so this costs no extra decode; then - the middle of the nearest pause (a detected silence) when no motion was measured; then
- the even point. A move that would leave a part shorter than 0.5 s falls back to the even point.
Sub-shots are ordinary shots with consecutive s{n} ids — the editor keeps part of a long shot by
keeping only some of them, still never naming a time — and carry Part ("2/3") in the artifact and
part in the view (absent on every unsplit shot). Splitting happens before ids are assigned and
before audio levels, captions and duplicate grouping, so each part gets its own measurements.
Author-declared scenes are never split. The VideoStoryEditor prompt (both copies) gained a short
"Parts of a long shot" section.
Output format
VideoCompileStepConfig.OutputFormat picks the canvas the program is encoded at, and OutputFit
how each clip is fitted into it. Source (default) keeps the clips' own size: a single clip is
encoded at its own resolution exactly as before, and several clips share the canvas of the
largest kept source (pixel area, ties to the clip kept first) — never merely the first clip's,
which used to downscale 1080p b-roll to whatever generated clip opened the edit. The fixed formats
are Vertical 1080x1920 (Reels, Shorts, TikTok), Square 1080x1080, Landscape 1920x1080 and
Portrait 1080x1350 (4:5 feed posts). A run can override OutputFormat on every compile step at
once — the quick format switch, which also turns the default Crop fit into Subject for a
portrait or square target — and the analyze view then tells agents each clip's orientation and the
target; see editing-styles.md.
OutputFit | What a clip of a different aspect gets |
|---|---|
Crop (default) | scale=W:H:force_original_aspect_ratio=increase,crop=W:H — fills the frame; a 16:9 shot in 9:16 keeps its middle third |
Pad | scale=W:H:force_original_aspect_ratio=decrease,pad=W:H — the whole shot, filled per CanvasFill (black bars, or a blurred copy of the shot) |
Subject | the Crop scale, with the crop window placed on the clip's subject instead of its middle — see Reframing to the subject |
Both encode paths fit. The single-source select path appends the crop/pad suffix to its cut stage,
before the grade, inserts, overlays and captions, so everything painted afterwards works in
canvas pixels. A single clip at Pad + CanvasFill = Blur needs a split/blur/overlay graph per
clip, so that one combination routes to the segmented path, which builds it. On the segmented path a
fixed format normalizes every span (fitAllSourcesToCanvas), even when there is only one clip.
Geometry follows the picture through one mapping, SourceCanvasFit.For(source, canvas, cover):
the contain-fit arithmetic pad uses, or with cover the scale-to-cover arithmetic crop uses,
whose offsets go negative because the picture overflows the canvas. Tracked-insert corners, censor
masks and overlay placement bands (MapRectThroughFit, clamped to the frame) are all mapped
through it; a placement band that a crop removes entirely is dropped as placement_cropped_out.
A fixed format with Mode = StreamCopy is the hard config error OUTPUT_FORMAT_REQUIRES_REENCODE
— resizing has no stream-copy equivalent. Every compile reports what it encoded:
"output": { "format": "Vertical", "width": 1080, "height": 1920, "fit": "Crop" },
"length": { "targetSec": 20, "actualSec": 10.0, "met": false }
output.fit is the configured fit at a fixed format, Pad for several clips on a source-sized
canvas, and None for a single clip at Source. length is present only when
TargetDurationSec > 0; met is false when the real length is more than 20% away from the target.
The compile never invents footage to meet a target — a short source makes a short video, and the
customer is told so (target_length_missed, see Compile warnings).
Target length
A cut that runs LONG is brought down to TargetDurationSec before the encode (and before
cutting to the beat, which only shortens further) by TargetLengthPlanner:
span tails are shortened under one common ceiling, longest spans first, so every kept span keeps its
opening and short beats are untouched. No span goes below the narration anchored in it (the same
floors beat sync uses) or 1 s, and the LAST span — the end card or closing beat — is never touched.
When those floors make the target unreachable the cut comes as close as they allow. length.fitted
reports editedSec (the editor's length), trimmedSec and trimmedSpans; it is absent when nothing
was trimmed. A POV teaser asked to run 15 s came out at 25 s while the length was only reported.
Reframing to the subject
A centre crop is wrong whenever the subject is not in the middle: reframing a 16:9 shot to 9:16 keeps
only x ∈ [0.34, 0.66], and a phone held at a third of the frame is cut out of its own shot.
OutputFit = Subject keeps the Crop scale but places each clip's window on what the clip is
about. The position is chosen per source clip, from what the analysis measured inside the parts of
that clip the edit KEEPS (CropFocusResolver), in this order:
- A tracked screen — every insert/object-track keyframe of that source inside a kept window
(
basis: "tracked"). Measured every frame and free, and on a tracked-insert reel it is exactly the thing the shot is about. - A located subject —
VideoAnalyzeStepConfig.LocateSubjectsasks theVisionprovider for one box per shot on the shot's keyframe (the sameIObjectLocatorobject tracking uses: integer 0..1000 coordinates, robust JSON repair, first box wins), then re-asks on a zoomed crop around it exactly as the object tracker refines its seeds — a full-frame box is loose (a laptop came back 0.39 of the frame wide around a 0.29-wide screen), and a loose box is what tips a clip into padding. Two calls per shot; a failed refinement keeps the coarse box. Stored asVideoAnalysisShot.Subjectand weighted by how long each shot is kept (basis: "subject").SubjectLabelnames what to look for ("the phone", "the presenter"); null asks for the main subject.MaxSubjectShots(40) caps the calls, and the step'sVisionTimeoutSecondsbounds them. A failed or empty lookup leaves the shot without a subject; it never fails the step. The box is in the artifact only — never in the view, never offered to an agent, and it can only ever move a crop. - Otherwise the centre (
basis: "centre"), which is byte-identical toCrop.
The window is static per clip. Along each axis it takes the position that keeps the most weighted subject inside it, and among equally good positions the one nearest the weighted centre: one subject that fits is centred in the window; two subjects too far apart to both fit leave the one on screen longer whole, rather than halving both (a weighted centre) or showing neither (a union centre).
When that window would keep less than 87% (CropFocusResolver.MinKeptFraction) of the clip's MAIN
object — the track or located subject on screen longest — no crop can honour "keep the subject in
shot", so that one clip is shown whole instead: padded per CanvasFill, exactly like
OutputFit = Pad (CropFocus.Pad, reported as "fit": "Pad" on its output.focus entry). An end
card whose headline is wider than a 9:16 window, or a monitor filling a 16:9 shot, keeps every word
and edge; a phone or a laptop screen that fits is cropped. Secondary subjects never trigger it — a
second box too far away to share the window is simply left out. A grid-saliency estimate (detail + motion on the 32x18 analysis grid) was measured first
and rejected: it locked onto sharp background clutter, placing the woman's phone shot at 0.56
against a true 0.33.
The crop is crop=W:H:'max(0,min(iw-ow,fx*iw-ow/2))':'max(0,min(ih-oh,fy*ih-oh/2))' after the
cover scale, on both encode paths (CoverCropFilter), and SourceCanvasFit.For(..., focus) mirrors
the same clamp so tracked inserts, censor masks and overlay bands stay registered to the picture
(verified pixel-identical against an explicit crop=W:H:x:0). A seam-bridge clip is always centred.
The compile reports where each clip's window went:
"output": { "format": "Vertical", "fit": "Subject",
"focus": [ { "source": 0, "x": 0.33, "y": 0.5, "basis": "subject" } ] }
Program loudness
VideoCompileStepConfig.LoudnessTarget (ProgramLoudness: Off default and byte-identical,
Social −14 LUFS / −1.5 dBTP, Web −16 LUFS / −1.5 dBTP, Broadcast −23 LUFS / −2 dBTP — EBU R128)
brings the finished program to the level the platform it is posted on plays back at. A narrated
launch reel measured −22.5 LUFS, about 8 dB quieter than everything around it in a feed.
It runs after the encode, in two ffmpeg passes over the encoded file (LoudnessNormalizer): a
loudnorm measuring pass, then an applying pass fed every measured value with linear=true — one
gain for the whole program, so the balance between narration, music and effects is unchanged; when
that gain would breach the true-peak ceiling, loudnorm itself switches to its dynamic mode. Only
the audio is re-encoded (AAC 192k, 48 kHz); the video stream is copied, so the picture is
bit-identical. The ceilings sit at −1.5 dBTP rather than −1 because the AAC encode overshoots: the
same reel normalized to a −1 ceiling measured −0.6 dBTP afterwards. A peak limiter (alimiter, half a dB under the
ceiling, lookahead compensated) follows loudnorm: when one linear gain cannot reach the target
without breaching the ceiling, loudnorm falls back to its dynamic mode, which does not hold the
ceiling — a narrated story came out at +0.5 dBTP without it and −1.8 dBTP with it (−14.3 LUFS).
The result is then measured again: the AAC encode after the limiter can still overshoot on hard
transients (a hip-hop reel measured −0.14 dBTP against −1.5). When the measured true peak is more than
0.2 dB over the ceiling, the pass is redone once from the original encode with the limiter lowered by
the overshoot plus 0.3 dB — one AAC generation, loudness unchanged (that reel: −3.4 dBTP, −14.3
LUFS). The node reports outputTruePeak and, when the redo ran, peakCorrectedDb.
It never fails the compile. The loudness node reports the target, the measured loudness and true
peak, the gain applied, or why nothing was done (no_audio_in_program, audio_stream_copied,
not_measurable for silence, measure_failed, apply_failed, normalize_failed).
"loudness": { "target": "Social", "targetLufs": -14, "targetTruePeak": -1.5, "applied": true,
"measuredLufs": -22.8, "measuredTruePeak": -3.8, "gainDb": 8.8 }
Program audio on silent footage
Stock and generated b-roll routinely has no audio stream. Before this fix a compile over such
footage with narration (or sound effects) and no music dropped audio altogether: the step reported
voiceover.applied: true, captions were burned in, and the MP4 had no sound track, because the
narration had nothing to be mixed into.
Now, whenever narration or sound effects are placed and neither the source nor a music bed provides a base, both encode paths synthesize one:
- Select path (one clip). The base is
anullsrc=r=48000:cl=stereo:d={program length}[abase](BuildSilentBaseBranch) — the same 48 kHz stereo every mix branch is normalized to, ending at exactly the program's length, so eachamix ... duration=firststage stays pinned to the picture. - Segmented path (several clips, transitions, generated clips). Every span gets the
matching-duration
anullsrcbranch an audio-less span already gets in a mixed edit, so the concat, crossfades and cold-open/end-card concat all carry an audio stream end to end.
With music, the music branch is the base exactly as before, and a compile with nothing to mix still
has no audio track; both filtergraphs are byte-identical to before the fix. The step output's
audio node keeps applied/reason describing the source dialogue and adds
silentBase: true, programAudio: true when a base was synthesized. Sound effects no longer need
dialogue or music to be heard (the old no_base_audio drop is gone).
Compile warnings
A render can complete and still not be what the customer asked for. The compile reports those cases
in a top-level warnings array (the shared C1 shape the execution page and the renders tab
display), present only when non-empty so a clean compile's output keeps its exact shape:
"warnings": [ { "code": "narration_mostly_dropped",
"message": "Only 1 of 8 narration lines made it into the video, so most of the narration is missing. ..." } ]
| Code | When |
|---|---|
narration_dropped | Some declared narration lines (voiceover.lineCount - appliedLineCount) are not in the render, but at most half |
narration_mostly_dropped | More than half are missing. The ReviewLoop also caps the review score at 4 for this shape (ReviewLoopStepExecutor.ApplyNarrationMostlyDroppedCap), so a review loop iterates instead of passing a mostly mute narrated video — see voiceover.md |
no_program_audio | Narration, music or effects were placed but the program has no audio track. Defensive — the silent base makes it unreachable |
music_no_tracks | EnableMusic is on, no fixed MusicTrackProjectFileId, and the analyze step offered no music candidates — the project has no audio files to choose from; the message tells the customer to upload one |
sfx_no_clips | EnableSfx is on, nothing was applied, and the analyze step offered no SFX candidates |
target_length_missed | TargetDurationSec > 0 and the real length is more than 20% away from it |
Output file names. With RegisterProjectFile on and OutputFileName left at edited.mp4, the
render is registered as "{workflow name} - run {n}.mp4", where n counts this workflow's
executions started up to this one, plus (revision {k}) when a review loop re-compiled. The name is
reduced to ASCII letters, digits, spaces and ( ) - _ . , (accents are folded, anything else becomes
a space — it travels into object-store metadata, which must be ASCII) and capped at 120 characters.
When no workflow name can be found the render keeps edited.mp4; a custom OutputFileName is
always kept. Storage keys are unchanged (they are id-based).
One registered render per run
With RegisterProjectFile on, every compile used to add another project-file row — every review-loop
pass, every retry, and every run of a workflow whose OutputFileName is fixed. One customer project
ended up with ~25 identical "20s narrated timeline teaser - run 2.mp4" entries in its Files tab.
After registering its render, the compile step now removes the earlier rows THIS workflow step
registered that the new render supersedes (VideoCompileStepExecutor.RetireSupersededRegistrationsAsync):
- every earlier render of the same step in the same execution (review-loop passes, retries), whatever it was called — the Files tab keeps the run's latest render;
- earlier renders of the same step from other executions only when they were registered under
the very same file name (a fixed custom
OutputFileName). Renders named"… - run {n}.mp4"are distinct per run and are kept: that is the per-run versioning.
A row counts as "this step's render" only when a WorkflowStepResult of this same WorkflowStep
names its storage key as OutputStorageKey — never by name alone — so a file a user uploaded, or
another step's render, is never touched. Only the rows are removed; the stored videos are not,
and every earlier render stays reachable from its own run's outputs (the Renders tab and the run
page read step results, not project files). The cleanup is best-effort: a database error leaves
the Files list as it was and the compile still succeeds.
When it removed anything, the output summary carries:
"registration": { "fileName": "Spring teaser - run 2.mp4", "projectFileId": "…", "replacedEarlierRenders": 3 }
(absent otherwise, so a first registration's output keeps its exact shape). A workflow that pointed
a later step at an earlier render's project-file id by hand will no longer find it; point it at the
latest render, or at the step's output (StepOutput).
Takes: the latest render first
Every run, review-loop pass and retry renders again, and the Renders tab and the Files view used to list every one of those renders side by side. They are now treated as takes of one video.
What a take is. Every render (a WorkflowStepResult with an OutputStorageKey under the
project's outputFiles/) of the same workflow step is a take of that step's video. The group key
is {workflowDefinitionId:N}:{stepOrder}:{stepType} (so a workflow re-saved with new step rows keeps
its takes together), or step:{stepId:N} when the workflow or step is gone
(OutputsController.TakeGroupKey).
The latest take is the newest render of a run that has ENDED. A render of a run still queued or running is never the latest take — it may still be replaced — and is never deleted. A render that is not the last one its own run produced for that step is intermediate (a review-loop pass or a retry the run then replaced).
API (/api/v1/projects/{projectId}/outputs, owner-scoped like every project endpoint: unknown
project 404, another owner's 403):
| Method | Path | What it does |
|---|---|---|
GET | /outputs | Every render, newest first, now also carrying takeGroupKey, takeIndex (1 = latest), takeCount, isLatestTake, isIntermediate, isPinned — appended to OutputVideoResponse, every older field unchanged |
POST | /outputs/{stepResultId}/pin | { "pinned": true } pins a take (WorkflowStepResult.IsPinned, column is_pinned, engine migration AddStepResultIsPinned); a pinned take is kept by every cleanup |
DELETE | /outputs/{stepResultId} | Deletes one earlier take; 409 for the latest take, a pinned take, or a take whose run is still going |
POST | /outputs/prune | { "takeGroupKey": "…", "keepLatest": 1 } ("Delete earlier takes") deletes every take of that video except the newest keepLatest (at least 1), the pinned ones and those of runs still going; returns { deleted, kept } |
Deleting a take removes its stored object — only when no surviving take still uses the same key
(a step-cache hit reuses an earlier run's object) — removes its project-file row if it was saved to
Files, and clears the step result's OutputStorageKey. The step result itself stays, so the run's
history (its output text, the EDL, the timing) is intact; the run page simply no longer shows a
video for it.
Web. The project's Renders tab and the Overview's "Latest renders" show each video once, by its
latest take (web/lib/outputs/takes.ts groupTakes, which falls back to workflow + step id against
an older server). Under the player, the video being watched lists its earlier takes collapsed under
"N earlier takes" (components/outputs/EarlierTakes.tsx): each playable and downloadable, marked
"In-progress take" (intermediate) or "Run still going", pinnable, and deletable after a
confirmation; "Delete earlier takes" asks once and calls prune. The Files tab hides the
project-file rows of renders that are not the latest take of their video
(earlierTakeStorageKeys — a key the latest take still plays from is never hidden) behind
"Show N earlier takes of your videos". There is no automatic keep-the-last-N setting: prune with
keepLatest is the building block for one.
See also One registered render per run, which stops the compile step from registering a Files entry per review-loop pass in the first place.
Scene/visual analysis (Phase 1)
Phase 1 adds deterministic (no LLM, no new external dependency) visual and audio descriptors per
shot on top of the shipped VideoAnalyze step, which until now only knew about cut points
(silence gaps, shot-change timestamps) and speech. It is purely additive: nothing existing changes
behavior when the new config defaults are used, other than the artifact and bounded view gaining
new optional fields (VideoAnalysisArtifact.Version moves to 2, but a Version: 1 artifact
still deserializes unchanged — every new field is optional/nullable/default-valued and appended,
never inserted or reordered).
The technique: one low-res raw-frame grid pass
A single new ffmpeg invocation decimates and downscales the source video to a raw RGB pixel grid
file (FfmpegFrameGridSampler, FfmpegArgvBuilder.BuildGridSampleArgs):
ffmpeg -nostdin -hide_banner -y -loglevel error -protocol_whitelist file
-i {input}
-an -sn
-vf fps={sampleFps},scale={gw}:{gh}:flags=area,format=rgb24
-f rawvideo -pix_fmt rgb24
{scratch}/grid.rgb
flags=area is load-bearing — it's a true box-average downscale, so each output pixel is the
exact mean of its source block, which is what makes per-region statistics meaningful. Default grid
is 32x18 at 2.0 fps; MaxVisualSampleFrames (default 4000) clamps the effective fps downward
for very long videos: effectiveFps = min(VisualSampleFps, MaxVisualSampleFrames / durationSec).
Sample time of raw frame i is exactly i / effectiveFps (the fps filter emits CFR from t=0),
so no timestamp parsing is needed anywhere downstream — just byte-offset math
(frameSizeBytes = gridWidth * gridHeight * 3).
FrameGridAnalyzer (WorkflowEngine/Services/Video/FrameGridAnalyzer.cs) is a pure,
unit-testable static class that derives everything below from this one grid buffer — no ffmpeg
stderr scraping for any of it:
- Motion — mean absolute luma delta between consecutive sampled frames within a shot,
normalized 0..1:
MotionMean/MotionPeak/MotionStdDev, bucketed into aMotionClass(Static/Subtle/Moderate/Dynamic). - Camera move — a 1-D SAD (sum-of-absolute-differences) integer pixel-shift search
(
dx/dyin[-4, 4]) between consecutive frames' column-sum and row-sum luma profiles. A consistent same-sign shift ⇒Pan/Tilt; a high-variance alternating-sign shift ⇒Handheld; more motion near the frame center than the border ⇒Zoom; near-zero shift and near-zero motion ⇒Static; otherwiseUnknown. Reported asCameraMove+CameraMoveConfidence(0..1) — this is a documented heuristic, not ground truth, and is always paired with its confidence. - Still windows / head-tail motion — runs where per-frame motion stays below
StillMotionThresholdfor at leastMinStillWindowMs(capped toMaxStillWindowsPerShot, longest kept), plusHeadMotion/TailMotion(mean motion over the shot's first/last 250ms) — "will a cut here land mid-motion?" - Exposure/color — Rec.709 luma per pixel, shot-averaged into
BrightnessMean/BrightnessStdDev(temporal flicker),ContrastRms(intra-frame luma std-dev, shot-averaged),ClippedHighlightRatio/CrushedBlackRatio,SaturationMean(HSV).DominantColors: the top 3 bins of a 64-bin (4 levels/channel) RGB histogram, each as a hex color + population share. - Regions / safe zones — the grid's 3x3 spatial cells (
R0..R8) plus three named overlay-candidate bands (LowerThird,UpperThird,CenterBand). Per region:LumaMean,LumaStdDev(clutter proxy),TemporalMotion,TextColor(Light/Dark, fromLumaMean), andSuitability(0..1) —0.5*clutterScore + 0.3*motionScore + 0.2*extremeScore, favoring an uncluttered (low luma std-dev), low-motion region whose brightness isn't at a 0/1 extreme.BestOverlayRegionon the shot is the highest-Suitabilityname among the three named bands. - Near-duplicate / best-take grouping — per shot, a signature (
FrameGridAnalyzer.ShotSignature): a z-normalized, time-averaged luma grid (LumaSig, 576 floats for 32x18) plus the L1-normalized 64-bin color histogram (ColorHist).Distance = 0.7*(1-cosine(LumaSig)) + 0.3*(0.5*L1(ColorHist)),Similarity = 1 - Distance. Groups form via single-linkage clustering over a sliding time window (DuplicateWindowShots, default 20 — only the previous N shots are compared, both faster and more correct since multi-take shots are temporally adjacent) atDuplicateSimilarityThreshold(default0.90). Each group gets an idd{n}; members are ranked by a heuristicTakeQuality(documented inFrameGridAnalyzer.ComputeTakeQuality: 35% inverse motion jitter, 25% audio level, 20% exposure, 10% duration, 10% neutral sharpness placeholder) — rank 0 isIsBestTake. A singleton shot gets no group (DuplicateGroupId = null). - Ken-Burns candidate —
KenBurnsCandidate = truewhen the shot isStatic,MotionMean < 0.02, duration>= 2.5s, and at least one named band has decentSuitability; identifies candidates only — no zoompan is ever applied (out of scope, as always). - Audio levels —
WavRmsSampler(pure,WorkflowEngine/Services/Video/WavRmsSampler.cs) windows the canonical 16kHz mono s16 WAVFfmpegAudioExtractoralready produces into 250ms RMS/peak windows (20*log10(rms/32768), floored at -96 dBFS instead of-Infinity); reuses the WAV already extracted for transcription, or extracts it fresh if transcription is off/degraded. Per shot:AudioRmsDbfs(energy-weighted mean of overlapping windows),AudioPeakDbfs, andSpeechRatio(from the already-detected silence spans — no new audio pass), bucketed into aLoudnessClass(Quiet/Normal/Loud). - Not implemented in Phase 1 (opt-in, default
false, lower priority than the core grid pipeline):DetectLetterbox/ActiveCropandDetectSharpness/Sharpness— the config fields and artifact columns exist (alwaysnull/false) so a future phase can fill them in without another schema migration;MetadataPrintOutputParser/CropDetectOutputParser(genericffmpeg ... metadata=print:file=-/cropdetectstderr parsers, mirroringSilenceDetectOutputParser/ShowinfoOutputParser's style) were likewise left unimplemented.
VisualDetail: degrade before drop
VideoAnalyzeStepExecutor.BuildBoundedView treats "how many shots/silences/segments the model
sees" as higher priority than "how much visual/audio detail each one carries". Before ever
dropping an offered item to fit MaxOutputChars, it tries the configured VisualDetail level,
then each lower level in turn — Full → Compact → None — re-serializing the same set of offered
items at each level. Only once None (which renders a shot exactly as it looked before Phase 1
— no v/a key at all) still doesn't fit does the pre-existing item-dropping loop run. This
guarantees meta.offeredIdCount can never be smaller than what a plain AnalyzeVisuals: false run
would produce at the same MaxOutputChars — richer per-shot data can only ever cost detail, never
cost coverage. meta.visual.detail (and the top-level meta.visualDetailApplied) records whichever
level actually got used.
None— identical to the pre-Phase-1 shot shape.Compact(default) —motion,move,cutIn/cutOut, the single longest still window,bright,contrast, the top 2 dominant colors,safe(best overlay region),dup/best,kenBurns(only when true), plusa: {rms, speech}when audio levels are available.Full— everythingCompacthas, plus the full region list, the full still-window list,motionStdDev/motionPeak/cameraConfidence, and all (up to 3) dominant colors.
Every visual/audio number in the view is rounded before serialization — seconds to 2 decimal
places, 0..1 scores to 0..100 integers (Round2/Score helpers) — since a raw double can
serialize as 15-17 characters of floating-point noise, which adds up fast across hundreds of
shots. The four base shot fields (id/startSec/endSec/durationSec) are deliberately left
unrounded at every detail level, matching the pre-Phase-1 output exactly.
View/artifact shape additions
{
"view": {
"shots": [{
"id": "s4", "startSec": 12.4, "endSec": 16.8, "durationSec": 4.4,
"v": {
"motion": 12, "move": "Pan", "cutIn": "still", "cutOut": "moving",
"still": [{ "startSec": 12.4, "endSec": 12.9 }],
"bright": 41, "contrast": 22, "colors": ["#2b3a4f", "#c9b48a"],
"safe": { "region": "LowerThird", "fit": 88, "text": "Light" },
"dup": "d2", "best": true, "kenBurns": true
},
"a": { "rms": -21, "speech": 82 }
}],
"pacing": { "meanShotSec": 4.1, "medianShotSec": 3.8, "cutsPerMinute": 14.6, "motionTimeline": [12, 30, 8], "timelineBinSec": 5.0 },
"duplicateGroups": [{ "id": "d2", "shotIds": ["s4", "s7"], "bestShotId": "s7", "similarity": 94 }]
},
"meta": {
"visual": { "applied": true, "degraded": false, "sampleFps": 2.0, "gridWidth": 32, "gridHeight": 18, "detail": "Compact" },
"audioLevels": { "applied": true }
}
}
Pacing (FrameGridAnalyzer.ComputePacing) is a whole-artifact summary — MeanShotSeconds/
MedianShotSeconds/CutsPerMinute from shot timing alone, plus a MotionTimeline (mean
MotionMean per TimelineBinSeconds-wide bin, empty when no shot has visual data). pacing/
duplicateGroups are only surfaced in the view at Compact/Full detail — never at None — so
that a budget-forced collapse to None (whether from AnalyzeVisuals: false or from a degraded
visual-analysis stage) is always byte-identical regardless of whether visual data merely got
suppressed by degradation vs. never computed at all.
Failure handling: degrade, never fail the step
Both the visual-analysis block (grid sampling + every FrameGridAnalyzer call) and the
audio-level block (WAV sampling) are wrapped in their own try/catch, exactly like the existing
transcription-degrade pattern: on any exception, Provenance.VisualAnalysisApplied/
AudioLevelsApplied become false, VisualAnalysisDegraded becomes true, a warning is logged,
and the step continues with shots/silences/transcript exactly as if that stage were configured
off. Nothing in Phase 1 can fail a VideoAnalyze step.
Vision captioning (Phase 2)
Optional vision-LLM shot captioning on top of Phase 1's deterministic descriptors: extract one
representative keyframe per selected shot, send it to a vision-capable chat model, and get back a
short structured scene description (subjects, action, setting, mood, shot scale, camera angle,
on-screen text, tags). Unlike every other stage in this document, this one is off by default
and makes real LLM calls — see Why Off, not Optional below.
InferenceProviderCapability.Vision
A third InferenceProvider capability, alongside Chat and Transcription. It reuses the
exact same chat-completions machinery Chat does — IChatClientFactory/ResolvedInferenceProvider
gained no new members — since a vision call is just an ordinary chat-completions call with an
image content part alongside the text prompt. It is still a separate capability (not folded into
Chat) so a vision-capable deployment (which may differ from the deployment an agent's chat
resolution uses) can be configured and defaulted independently: Vision participates in its own
"at most one default" bucket via the same composite (capability, is_default) partial unique
index Chat/Transcription already share — no schema migration was needed to add the third
value, since that index is generic over any capability value, not hardcoded to two.
IInferenceProviderResolver.ResolveVisionAsync(explicitProviderId, ct) mirrors
ResolveTranscriptionAsync's precedence and null-when-nothing-resolves contract exactly
(explicitProviderId if it resolves to an enabled Vision-capability row → the single enabled
IsDefault && Capability == Vision row → null), but returns ResolvedInferenceProvider? (not a
dedicated vision type) and resolves through the same ResolveFromProvider helper ResolveAsync
(chat) uses. Like transcription, there is deliberately no fallback to the legacy AzureOpenAI:*
config keys — silently sending an image to a deployment that may not support vision would fail
confusingly.
POST /api/v1/inference-providers/{id}/test and POST /api/v1/inference-providers/test gained a
Vision arm alongside the existing Transcription arm: it sends a trivial embedded 1x1 JPEG
through IChatClientFactory with a "reply with the word ok" prompt and treats any non-empty
response as success — mirroring the existing synthesized-silent-WAV transcription ping. The admin
UI (InferenceProviderForm) exposes Vision as a third Capability option; the per-agent
provider override picker (AgentInferenceProviderSelect) filters to Chat rows only (an
allowlist, not merely "not Transcription" — a denylist would have silently admitted Vision rows
here too), since that override feeds chat resolution only.
Why Off, not Optional
VideoAnalyzeStepConfig.Transcription defaults to Optional because it costs at most one ASR
network call per step. Captioning is structurally different: it can cost up to MaxCaptionedShots
(default 24) separate vision chat-completion calls — each carrying an image — per VideoAnalyze
step. If Vision defaulted to Optional, then the moment any admin configured a
Vision-capability default provider for some unrelated workflow that actually wants captioning,
every other existing or future VideoAnalyze step in the system would silently start making
real, billed vision calls with no config change of its own. Defaulting to Off keeps every step's
cost/latency unchanged unless its author explicitly opts in by setting Vision on that step.
Optional/Required otherwise carry the same degrade-vs-fail semantics Transcription does.
Keyframe selection
KeyframeSelector (WorkflowEngine/Services/Video/KeyframeSelector.cs) is pure — no ffmpeg, no
I/O:
ChooseKeyframeSec(shot)— the midpoint of the shot's longestStillWindow(Phase 1) when at least one exists (a calm moment makes a cleaner, less motion-blurred frame), else the shot's own midpoint. Falls back to the shot midpoint whenVisualisnull(visual analysis off, degraded, or simply no still windows).SelectShotsToCaption(shots, duplicateGroups, strategy, maxCaptionedShots, minCaptionShotSeconds)implements three strategies (VideoCaptionSelection), excluding any shot shorter thanminCaptionShotSeconds, and always returning ids in shot-chronological order regardless of selection order:PerDuplicateGroup(default) — captions each near-duplicate group's best-take shot first (Phase 1'sDuplicateGroups), so N takes of one setup cost one vision call, not N; fills any remaining budget with the longest not-yet-selected shots.LongestShots— simply the N longest eligible shots.EvenlySpaced— shots at roughly even index intervals across the whole shot list.
Keyframe extraction
FfmpegArgvBuilder.BuildKeyframeArgs(inputPath, outputJpgPath, atSec, maxWidth) — a new ffmpeg
invocation alongside Phase 1's grid-sample/audio-extract builders, following the exact same
IVideoToolRunner calling convention via a new IKeyframeExtractor/FfmpegKeyframeExtractor pair
(mirroring IAudioExtractor/FfmpegAudioExtractor):
ffmpeg -nostdin -hide_banner -y -loglevel error -protocol_whitelist file
-ss {atSec} -i {inputPath}
-frames:v 1
-vf scale='min({maxWidth},iw)':-2
-f image2 -c:v mjpeg -q:v 4
{outputJpgPath}
-ss before -i for fast input seeking (same convention as BuildExtractAudioArgs); the scale
expression never upscales and preserves aspect ratio. IKeyframeExtractor is deliberately its own
small interface (not folded into IFrameGridSampler) so Phase 3 (motion-graphics overlay
planning, applied at compile time) has a single established pattern to follow for its own new
ffmpeg operations.
The captioner
IShotCaptioner/VisionShotCaptioner (WorkflowEngine/Services/Video/) build one chat message
per shot — a text prompt plus the keyframe JPEG as a DataContent("image/jpeg") content part —
and call IChatClient.GetResponseAsync with ChatResponseFormat.ForJsonSchema<VideoShotCaption>(),
the exact same structured-output mechanism AgentStepExecutor/ReelBoltAgentBase use for agent
steps. IChatClientFactory is reused unchanged.
public sealed record VideoShotCaption(
string ShotId, string Summary, IReadOnlyList<string> Subjects,
string Action, string Setting, string Mood,
string ShotScale, string CameraAngle,
IReadOnlyList<string> OnScreenText, IReadOnlyList<string> Tags);
Critical safety property: the shot-id ↔ caption binding is never model-controlled. The model is
called once per shot; VideoAnalyzeStepExecutor — never the captioner — always overwrites the
returned VideoShotCaption.ShotId with the id it actually requested (ShotCaptionRequest.ShotId)
before attaching the caption to a shot. This mirrors the id-anchored discipline
VideoEditDecisionOutput/VideoCompileStepExecutor already use: never trust an identifier the
model echoes back for anything that matters. Covered by a dedicated test asserting the executor
ignores a deliberately-wrong model-returned ShotId.
Captioning retries up to 2 attempts per shot with a short linear backoff (mirrors
TranscribeWithRetryAsync's shape); a captioning failure on one shot never aborts captioning of
the rest — the executor's per-shot loop catches, counts it in meta.vision.failedShots, and moves
on.
Executor wiring
Captioning runs last among VideoAnalyze's analysis stages, strictly after every deterministic
stage (silence/shot detection, transcription, Phase 1 visual/audio analysis, near-duplicate
grouping) — so a vision failure can never put anything deterministic at risk:
Multi-source hoist (Phase 4 fix). Captioning is a step-level pass over every source's shots
combined, run once after all per-source deterministic analysis completes — not a per-source pass
run once per source. This matters for MaxCaptionedShots/VisionTimeoutSeconds: both are genuinely
step-wide budgets across every source clip, never silently reset per source. An earlier draft of
this feature ran captioning inside the per-source loop, which would have let a 2-source analysis
spend up to 2 × MaxCaptionedShots vision calls instead of the configured cap — caught and fixed
before merge; VideoMultiSourceTests pins the corrected step-wide behavior.
Vision == Off→ skipped entirely,meta.vision = {mode: "Off", applied: false, ...}. Every other field in the view/artifact is byte-identical to a pre-Phase-2 run — the single most important regression test in this phase, mirroring Phase 1's degrade-before-drop test's importance.- Else, resolve via
ResolveVisionAsync.null+Required⇒ fail withVISION_UNAVAILABLE.null+Optional⇒ degrade (meta.vision.degraded = true), continue with no captions. - Resolved ⇒ select shots via
KeyframeSelector, extract each keyframe to scratch, caption each with the 2-attempt retry. An aggregateVisionTimeoutSecondswall-clock budget covers the whole captioning pass (not per-shot) via a linked, timedCancellationTokenSource; exceeding it mid-pass stops captioning further shots, keeps whatever already succeeded, and setsmeta.vision.partial = true. - If zero captions were obtained after all that:
Required⇒ fail withVISION_FAILED;Optional⇒ degrade and continue with zero (or partial) captions.
View/artifact shape addition
A shot with a caption gains a "c" key in the bounded view, gated on VisualDetail exactly like
Phase 1's "v"/"a" keys — same degrade-before-drop discipline, never bypassed:
{
"view": {
"shots": [{
"id": "s4", "startSec": 12.4, "endSec": 16.8, "durationSec": 4.4,
"c": {
"summary": "A presenter gestures at a whiteboard while explaining a diagram.",
"scale": "Medium", "mood": "Focused", "tags": ["presenter", "whiteboard", "explaining"],
"subjects": ["presenter"], "action": "gesturing at a diagram",
"setting": "office whiteboard", "cameraAngle": "Eye level", "onScreenText": [],
"style": "flat ungraded log", "issues": ["soft focus"]
}
}]
},
"meta": {
"vision": { "mode": "Optional", "applied": true, "provider": "gpt-4o-mini-vision", "degraded": false, "partial": false, "captionedShots": 6, "failedShots": 0, "persistedKeyframes": 0 }
}
}
Compact detail shows summary/scale/mood/tags/style/issues (Phase 4 added style and
issues at Compact, matching Phase 1's own "cheap signals first" discipline); Full adds
subjects/action/setting/cameraAngle/onScreenText/timeOfDay/lighting/framing (Phase
4). A shot with no "c" key is normal — not selected for captioning, or captioning off/failed/
degraded — never a signal the shot is empty or unimportant; the VideoStoryEditor prompt says so
explicitly.
VideoAnalysisArtifact.Version stays at 2, not 3: Phase 2 appends exactly one more
optional/nullable field (VideoAnalysisShot.Caption) plus optional/default-valued fields on
VideoAnalysisProvenance, and a Version-2-without-captions artifact and a
Version-2-with-captions artifact are both valid under the identical shape — no consumer needs to
structurally distinguish them (a consumer that cares simply checks whether Caption is null).
Phase 4's five new VideoShotCaption fields (TimeOfDay, Lighting, VisualStyle, Framing,
TechnicalIssues) are additive the same way and do not bump the version either.
Prompt priming, contact sheets, and persisted keyframes (Phase 4)
Three changes to the captioning path itself, none of which touch the id-anchored/no-timestamp contract:
- Prompt priming. The vision prompt is primed with a short sentence of Phase 1's own
deterministic measurements for the shot being captioned (color temperature, tone curve,
saturation, camera move, letterbox/backlit flags) — words derived from measurements, never a
number the model could restate. The model is told to use this only to inform its reading, not to
treat it as ground truth it must repeat;
null(visual analysis off/degraded for that shot) omits the block entirely.ShotCaptionRequest.MeasuredContextcarries this;VisionShotCaptionerformats and injects it. - Contact-sheet keyframes (
KeyframesPerShot, default1).1extracts the same single mid-shot still as before Phase 4 — byte-identical.2or3instead extracts that many frames evenly spaced across the shot (KeyframeSelector.ChooseKeyframeSecs) and hstacks them into one contact-sheet JPEG (FfmpegArgvBuilder.BuildContactSheetArgs,IKeyframeExtractor.ExtractContactSheetAsync) — N cheap input seeks, never a single pass decoding the whole shot. Each pane is scaled toKeyframeMaxWidth / Nso the combined image and its vision-call token cost stay roughly flat. The prompt gains one extra sentence telling the model the image is a multi-frame contact sheet of one shot, not several shots. PersistKeyframesis now wired:trueuploads each captioned shot's keyframe JPEG (or contact sheet) to storage undervideo-analysis/{executionId}/step-{stepOrder}-keyframes/{shotId}.jpg, counted inmeta.vision.persistedKeyframes. A persist failure is logged and swallowed — it can never cost the caption itself, since observability is never allowed to be more load-bearing than the thing it observes. Default staysfalse(scratch-only, deleted with the rest of scratch space).
Also (judgment call 9): KeyframeSelector's longest-shots fill pass now round-robins its selection
across source clips (grouping candidates by SourceIndex, longest-first within each group, then
alternating groups) instead of picking the global longest shots regardless of source — so a
MaxCaptionedShots budget spent across several source clips isn't silently exhausted entirely on
one long clip. A single-source analysis produces exactly one round-robin "group", which is provably
identical to the pre-Phase-4 plain longest-first order — no behavior change for the (still
overwhelmingly common) single-source case.
Builder UI for Phase 2 and Phase 4
The per-step VideoAnalyze config panel now has a "Vision captioning" section covering Vision,
VisionProviderId (filtered to Vision-capability provider rows by the same allowlist discipline
the per-agent chat override uses — never "not Transcription"), CaptionSelection,
MaxCaptionedShots, KeyframesPerShot, KeyframeMaxWidth and PersistKeyframes, alongside the
three free Phase 4 switches — see
Semantic visual dimensions (Phase 4). The remaining
fine-tuning fields (MinCaptionShotSeconds, VisionTimeoutSeconds, MaxCaptionChars,
MaxSharpnessShots, the look-grouping thresholds) stay settable via the raw step JSON only —
they are cost/quality trim, not pipeline shape.
Transcription (ASR)
Ships in full — not deferred — via the same InferenceProvider model chat completions already
use, distinguished by a new Capability column (Chat or Transcription). A single provider row
can't serve both roles: a Whisper deployment is a different deployment from a chat deployment, and
many OpenAI-compatible chat gateways have no /audio/transcriptions endpoint at all (only
whisper.cpp-server/faster-whisper-server/speaches/LiteLLM-style deployments do). Chat and
Transcription each have their own independent "at most one default" constraint (a composite
unique index on (capability, is_default)), and every resolution path — chat and transcription —
filters on its own Capability explicitly.
VideoAnalyzeStepConfig.Transcription has three modes:
Off— silence + scene + loudness only. Deterministic, zero external dependency, always available.Optional(default) — attempt ASR; if noTranscription-capability provider resolves, or the call fails after a small bounded retry, degrade cleanly toOffand recordmeta.transcription.degraded = truein the bounded view.Required— fail the step with a precise diagnostic if ASR is unavailable, rather than silently shipping an edit decision with no transcript context.
Long audio is chunked at silence-boundary-aligned points (never mid-word) by the pure,
independently-tested TranscriptChunkPlanner, and each chunk's word/segment timestamps are
offset by that chunk's absolute start before concatenation — the single most likely correctness
bug in this feature class, and the reason it has a dedicated multi-chunk offset test.
The admin UI for configuring providers (/admin/inference-providers) exposes Capability as a
field on create/edit; the per-agent chat-provider override picker filters Transcription rows
out, since that override only ever feeds chat resolution.
Transcripts from clips with no speech
Whisper-family recognisers fed music, room tone or silence routinely "hear" stock phrases — "Thank
you for watching.", "you", a subtitle credit. Before this gate a VideoAnalyze step offered such a
line to the story editor as real dialogue (a customer's b-roll edit chose its closing shot because of
one). TranscriptSpeechGate (WorkflowEngine/Services/Video/TranscriptSpeechGate.cs, pure) now
discards those segments per source clip, before t{n}/w{n} ids are assigned, so the id
namespaces stay contiguous and a dropped segment can never be offered, quoted to an agent, screened,
or resolved by VideoCompile. A segment is dropped when any of these holds:
no_speech— the recogniser says so itself: Whisper's own rule,no_speech_prob ≥ 0.6withavg_logprob < -1(or no log-prob reported), orno_speech_prob ≥ 0.9outright. The figures come fromverbose_json(TranscriptSegment.NoSpeechProb/AvgLogprob/CompressionRatio, optional members read by both transcription clients; a backend that omits them simply skips this rule).repetition—compression_ratio > 2.4, Whisper's looping-output signature.no_speech_energy— the step's own measurements: silence detection covers ≥ 80% of the segment's window, or no Phase 1 audio-level window under it rises aboveSilenceThresholdDb.stock_phrase— the whole segment is one of a short list of phrases recognisers invent on non-speech audio, and nothing vouches for it: the recogniser reportedno_speech_prob ≥ 0.2, or reported no confidence at all and the segment is the clip's only one. A confidently recognised "thank you" inside real dialogue is kept.
Words whose midpoint falls inside a dropped segment (and no kept one) go with it. A clip with no audio stream never reaches the gate: it is not transcribed at all, so it gets no transcript.
What was dropped is recorded, never silent: meta.transcription.droppedNoSpeech (present only when
non-zero, so a clean transcript's meta is unchanged) and
VideoAnalysisProvenance.TranscriptSegmentsDroppedNoSpeech in the artifact, summed across sources.
The VideoAnalyze revision tag in DeterministicStepRevisions moved with it, so cached analyses are
recomputed. Tests: TranscriptSpeechGateTests (each rule) and the "Transcript speech gate" block of
VideoAnalyzeStepExecutorTests (the hallucinated outro never reaches the view; t ids stay
contiguous; a no-audio clip is never sent to the recogniser).
Motion graphics (Phase 3)
Optional motion-graphics overlays — lower-thirds, titles, callouts — applied during
VideoCompile's encode, planned by a third built-in agent that reasons over overlay-safe-zone
"placement" candidates derived deterministically from Phase 1's per-shot region data, but never
emits a timestamp or pixel coordinate — the same structural discipline
VideoEditDecisionOutput/VideoStoryEditorAgent already established, extended to cover geometry
as well as time. Off by default (VideoAnalyzeStepConfig.EmitOverlayPlacements = false,
VideoCompileStepConfig.EnableGraphics = false) — both flags must be explicitly opted into, and
EnableGraphics = false leaves the compile path byte-identical to the pre-Phase-3 behavior.
This is the highest-risk phase of the feature: motion-graphics overlay TEXT is the first model-authored content in this feature to reach ffmpeg at all (every prior stage passes only opaque ids and workflow-author-supplied enum/allowlisted config). The text-sanitization and textfile-based discipline below is the core deliverable of this phase, not polish on top of it.
Placement candidates (deterministic, in VideoAnalyze)
OverlayPlacementBuilder (WorkflowEngine/Services/Video/OverlayPlacementBuilder.cs) is a pure,
unit-tested static class deriving overlay-placement candidates from shots that already carry Phase
1 Visual/Regions data (nothing is computed if visual analysis is off or degraded). For each
shot: rank its three NAMED overlay-safe-zone bands — LowerThird/UpperThird/CenterBand (never
the 3x3 R0..R8 grid cells, which exist for other diagnostics, not placement) — by
Suitability descending, take up to MaxPlacementsPerShot, and resolve a time window:
LongestStillWindow— the shot's longest Phase 1 still window, if one exists (clamped to the shot's own bounds).ShotMiddle— else, a ~2.5s window centered on the shot's midpoint, clamped to the shot's own bounds.ShotStart— else (a degenerateShotMiddlewindow, e.g. an extremely short shot), a window anchored at the shot's start, clamped to the shot's own bounds.
The whole artifact is capped at MaxPlacements, dropping the lowest-suitability candidates first;
survivors are then re-sorted into deterministic generation order (shot order, then per-shot rank)
and assigned SEQUENTIAL, globally unique ids p0, p1, … — a single counter across the whole
artifact, not per-shot.
Critical id-isolation requirement: p{n} placement ids are a SEPARATE namespace from
s{n}/g{n}/t{n} cut-anchor ids. VideoAnalysisArtifact.OfferedPlacementIds is a SEPARATE list
from OfferedIds, gated on the exact same "degrade before drop" VisualDetail discipline Phase 1
established: view.placements is shown as one atomic array at Compact/Full detail and omitted
entirely at None — so a placement id is only ever "offered" when the whole array survived to
whatever detail level the view actually settled on. VideoCompileStepExecutor.BuildIdTimeIndex
(which resolves Keep span ids to cut times) deliberately does NOT include placements — a code
comment there explains why — and a dedicated test proves a Keep span naming a p0 id fails
UNKNOWN_ID exactly like any other id that index does not contain.
view.placements entries are deliberately small: {id, shotId, region, startSec, endSec, fit, text} (plus src in a multi-source run) — a 2dp-rounded source-timeline window, a 0-100
suitability score, and a "Light"/"Dark" text-color hint, never a rect (geometry is resolved
server-side only, at compile time, from the full artifact).
When the MotionGraphicsPlanner step runs with agentInputContextMode: FullWorkflow (the only
mode that gives it the story editor's decision at all), each placement also gets an inEdit: bool
— whether that candidate's own moment survives the cut, computed by
MotionGraphicsPlacementAnnotator at prompt-assembly time in AgentStepExecutor (display-only:
nothing is filtered, an inEdit: false candidate stays fully choosable) by resolving the kept
spans through the exact same VideoCompileStepExecutor.MapSourceWindowToOutput that later drops
an overlay with reason cut_away — so the planner sees the survival verdict BEFORE it picks,
instead of discovering it only after the fact at compile time. Absent entirely under
PreviousStepOnly/other context modes, since there's no decision yet to check against.
Sentence-boundary and punctuation-reliability signals. Every view.segments transcript
segment also carries endsSentence: bool — whether that segment's own text (its full,
untruncated form, not the MaxSegmentTextChars-shortened copy the view shows) ends in terminal
sentence punctuation. This is the authoritative signal AgentType.VideoStoryEditor is told to use
instead of eyeballing the text field itself for a mid-sentence cut. But punctuation output is
only as trustworthy as the ASR backend that produced it — real deployments have been observed
transcribing some clips with almost no terminal punctuation at all even though every answer is a
complete thought — so meta.transcription.punctuated carries one entry per source clip,
{src, segments, punctuatedSegments, ratio, reliable} (reliable is ratio >= 0.5), and the
prompt is told to fall back to judging sentence-completeness semantically from the segment text
whenever the relevant source's entry says reliable: false. The same ratio/reliable pair
(recomputed for the specific source the compile step's sentenceCheck is about, not the whole
run) is surfaced to AgentType.VideoReviewAgent too, whose otherwise-hard "ends mid-sentence"
score cap is downgraded to a soft, non-blocking observation when that source's signal is
unreliable — a punctuation-starved ASR result must not by itself force a low score on an
otherwise-good edit.
startSec/endSec are input to the planner, not output from it, and that distinction is the
whole no-timestamp rule: MotionGraphicsPlanOutput still has no time-bearing property whatsoever
(MotionGraphicsPlanOutputInvariantTests), and VideoCompileStepExecutor still resolves an
overlay's real window from the full artifact's own unrounded
VideoAnalysisPlacement.StartSec/EndSec — never from the rounded numbers the view showed. The
precedent is view.segments, which has exposed exactly these two field names to
AgentType.VideoStoryEditor since the feature's first phase.
They are there because without them the view was undecidable. MaxTimeSlicesPerRegion splits one
long shot's region into several candidate sub-windows, so on a static shot (a talking-head
interview, say) several placements serialized byte-identically except for their id:
{"id":"p6","shotId":"s1","region":"LowerThird","fit":69,"text":"Light"}
{"id":"p7","shotId":"s1","region":"LowerThird","fit":69,"text":"Light"}
The planner had no information on which to prefer one over another and defaulted to the first of
each identical run — an arbitrary choice dressed up as an editorial one. The time window is the
only thing that distinguishes them, and it is also what lets the planner line an overlay up with
the view.segments entry (same units, same source timeline) whose content the overlay is about.
AgentType.MotionGraphicsPlanner and MotionGraphicsPlanOutput
An ordinary StepType.Agent step — no new step type, exactly the precedent
AgentType.VideoStoryEditor set. Given the story editor's decision (or the same bounded view) plus
view.placements, it decides zero or more overlays:
public class MotionGraphicsOverlay
{
public string PlacementId { get; set; } = ""; // must be in OfferedPlacementIds
public string Kind { get; set; } = ""; // LowerThird | Title | Callout | Tag
public string Text { get; set; } = "";
public string Subtext { get; set; } = "";
public string Duration { get; set; } = ""; // Short | Medium | Hold — never a number
public string Emphasis { get; set; } = ""; // Subtle | Normal | Strong
public string Reason { get; set; } = "";
}
public class MotionGraphicsPlanOutput
{
public List<MotionGraphicsOverlay> Overlays { get; set; } = new();
public string PlanRationale { get; set; } = "";
}
Every property on both types is a string/List<string-bearing-type> — the SAME rushcut
invariant VideoEditDecisionOutput established, extended to also forbid a pixel coordinate.
MotionGraphicsPlanOutputInvariantTests mirrors VideoEditDecisionOutputInvariantTests's
reflection approach exactly. Duration/Emphasis/Kind are enum WORDS the model chooses from a
closed vocabulary described in its prompt — VideoCompileStepExecutor alone resolves Duration
to milliseconds (OverlayShortMs/OverlayMediumMs/OverlayHoldMs, defaulting to Medium on an
unrecognized value) and OverlayFontSizePct to an actual pixel font size, which Emphasis
(Subtle/Normal/Strong) then nudges up or down by a fixed multiplier (0.8x/1x/1.25x, still
clamped to the same valid [2, 12] percentage range) in DrawtextFilterBuilder.ComputeFontSize —
a deliberately small effect, not a whole per-emphasis styling system. Kind
(LowerThird/Title/Callout/Tag) is currently descriptive/reserved only — nothing reads
it downstream, every kind renders identically, since the placement's own region
(LowerThird/UpperThird/CenterBand) already resolves the overlay's geometry and letting Kind also
influence position would create two disagreeing sources of geometry for the same overlay. See the
doc comment on MotionGraphicsOverlay.Kind in OutputSchemas.cs. The
MotionGraphicsPlannerAgent class's fallback prompt and the seeded built-in AgentDefinition row
in DatabaseSeeder are kept verbatim-identical, enforced by the second [Fact] in
VideoStoryEditorPromptConsistencyTests.cs
(MotionGraphicsPlanner_fallback_prompt_matches_the_seeded_built_in_agent_prompt_verbatim,
mirroring the first fact's reflection approach for VideoStoryEditorAgent). Tool access is now the
same full sandbox+Remotion+render pipeline AuthorAgent gets, minus WriteProjectFile
(AgentToolProvider) — widened from the minimal read-only scope VideoStoryEditor gets, since this
agent can optionally back an overlay with a real, rendered Remotion asset (see RenderedAssetStorageKey
below) rather than only plain drawtext/drawbox text. This is the only other agent besides AuthorAgent
granted RenderVideoAndUploadToStorage, and it is a conscious tradeoff: this agent's prompt includes
analysis-view content derived from the source video itself (on-screen text the Phase 2 vision model
read, ASR transcript text), so it is the first agent in this feature with code-execution tools whose
prompt is not limited to user-selected project files. The sandbox's own containment (read-only rootfs,
no network egress by default, no Docker-socket access — see Security)
is what bounds the blast radius of a successful prompt injection here: worst case is sandbox-contained
code execution, not host compromise. See CLAUDE.md's "Agent Types (enum)" section for the same note.
Rendered-asset overlays
An overlay can optionally carry MotionGraphicsOverlay.RenderedAssetStorageKey — the S3 storage key
of a transparent-background motion-graphics asset the MotionGraphicsPlanner agent produced ITSELF,
by actually calling RenderVideoAndUploadToStorage (a real tool call performing a real Remotion
render and a real upload), rather than an id merely echoed back from a set the model was shown. Still
not trusted blindly: VideoCompileStepExecutor validates the key matches the exact
projects/{projectId}/outputFiles/{executionId}/... prefix RenderVideoAndUploadToStorage itself
constructs for the CURRENT execution, before downloading or compositing anything at that key.
A rendered asset is the DEFAULT, not an upgrade. The tool grant that gives
MotionGraphicsPlanner/MotionGraphicsDirector the full sandbox+Remotion+render set exists for
this and nothing else, and the fallback a plain-text overlay reaches is drawtext over
drawbox — type on a translucent rectangle, visibly cheaper than anything the agent could
design. Both prompts originally framed rendering as "preferred when the moment deserves it" and
plain text as "the simple fallback"; run for real, the graphics room rendered one asset out of
three overlays on one pass and none out of three on the next, with no failed attempt in either —
it simply took the cheaper branch. The prompts, the room charter and the CopyArtist seat persona
now all state the same thing: every overlay is a designed render unless that specific moment has a
specific reason for unadorned type, stated in the overlay's own reason. A render that fails
after its retry budget may still fall back to text for that ONE overlay.
When present, OverlayAssetFilterBuilder (WorkflowEngine/Services/Video/OverlayAssetFilterBuilder.cs)
composites the asset via ffmpeg's overlay filter — added as an extra input, stretch-scaled to the
same compact ACCENT box geometry DrawtextFilterBuilder computes for a text overlay at the same
placement (DrawtextFilterBuilder.ComputeAccentBoxPixels, shared by both builders), and time-shifted
with setpts so the asset's own frame 0 lands at the overlay's actual on-screen start time on the
OUTPUT timeline. eof_action=pass means once the asset's content runs out the overlay simply stops
contributing (reverts to the plain cut) rather than freezing on its last frame for a longer "Hold"
window. Text/Subtext are ignored for an overlay that carries a rendered asset — an overlay is one
or the other, never both; a workflow author combines a rendered graphic with separate caption text by
authoring two overlays at different placements. Empty/absent (the default) leaves an overlay a plain
text overlay exactly as before this field existed — additive, not a replacement.
Alpha channel is validated, never trusted. Two independent guards close the "a real, transparent Remotion render still comes out as a solid, opaque block over the edited video" failure mode (observed in production as a black square with giant drop-shadow silhouettes stamped into it):
- Server-side validation — after downloading and ffprobing the asset,
VideoCompileStepExecutorchecks the probedpix_fmtagainstAlphaPixelFormats.HasAlpha(an allowlist of alpha-carrying formats —yuva420p/yuva444p10le/rgba/argb/gbrap/... — same allowlist discipline as the codec/color allowlists elsewhere in this file). An asset with no real alpha plane — e.g. the agent ignored the documented--pixel-format=yuva420p --codec=vp9render recipe and produced a plain H.264 file — is dropped for THIS overlay only (droppedOverlaysreasonasset_missing_alpha_channel), the same per-overlay degrade-not-fail discipline as a corrupt or unprobeable asset (asset_download_or_probe_failed), never composited opaque. - Filtergraph format-pinning — even when the source genuinely carries alpha,
OverlayAssetFilterBuilderprependsformat=rgbaas the FIRST operation on the overlay's own input, beforescale. This guards against a well-known libavfilter gotcha: format negotiation betweenscaleand the downstreamoverlayfilter can silently agree on a non-alpha common pixel format even when the decoded input has a real alpha plane, discarding transparency with no error of any kind. Pinning the format immediately after decode, before any negotiation happens, is what makes a genuinely transparent render actually composite transparently.
The source-to-output timeline mapping problem
A placement's window was resolved against the SOURCE video during VideoAnalyze. But drawtext's
enable=/alpha= expressions run against the ffmpeg filtergraph's OUTPUT timeline — the one the
existing select/setpts cut stage produces, which is shorter than the source and has every cut gap
removed entirely. Two internal, directly-unit-tested static functions on
VideoCompileStepExecutor solve this purely from the already-resolved, frame-quantized
ResolvedSpan list (their SnappedStart/SnappedEnd — the ACTUAL output-determining times, never
the pre-quantization requested ones):
internal static double? MapSourceToOutputSec(IReadOnlyList<ResolvedSpan> spans, double sourceSec);
internal static (double Start, double End)? MapSourceWindowToOutput(
IReadOnlyList<ResolvedSpan> spans, double startSec, double endSec);
Both walk spans accumulating output-timeline duration as they go. MapSourceToOutputSec returns
null when the second falls inside a cut gap (no corresponding output frame exists).
MapSourceWindowToOutput intersects a window with the kept spans and returns the output-timeline
window for the FIRST kept portion it overlaps (a single on-screen overlay cannot span a gap in the
output video) — null if the window never overlaps any kept span at all. Resolving a placement's
on-screen window intersects FIRST, truncates SECOND: the full candidate window
[placement.StartSec, shot.EndSec] (clamped to the OWNING SHOT's own bounds — looked up via
placement.ShotId against the full artifact's Shots, not merely the placement's own
already-narrow window) is mapped through MapSourceWindowToOutput to find where it survives the
cut, and only THEN is the model's chosen Duration applied, from wherever that surviving
intersection begins. Truncating to durationMs before intersecting (the original, buggy order)
silently dropped overlays whose full window overlapped a kept span by many seconds but whose first
durationMs alone did not — verified live: a placement with real window [0, 29.83] and a kept
span [4.8, 51.4] was tested only against [0, 3.0] (entirely cut) and wrongly dropped as
cut_away.
Text sanitization and the textfile=/expansion=none discipline
This is what makes the existing "not one model-originated character reaches an ffmpeg argv" claim (see Security) need a qualification, not a retraction. Overlay text is model-authored and does reach ffmpeg for the first time in this feature — but only as sanitized FILE CONTENT, never as argv or filter-string content:
OverlayTextSanitizer.Sanitize(raw, maxChars)(WorkflowEngine/Services/Video/) — an ALLOWLIST (never a denylist, which is only ever safe against characters someone thought of) of letters, digits, combining marks, spaces, and a small safe punctuation set. NFC-normalizes first; collapses ALL whitespace (including newlines/tabs — drawtext treats a raw newline as a forced line break) to single spaces before the allowlist strips anything, so a newline becomes a space rather than being silently deleted (which would wrongly glue two words together); truncates tomaxCharswithout splitting a grapheme cluster (StringInfo-based, never a blindstr[..n]); returns""for an empty/whitespace-only result, and the caller then drops that overlay/line entirely. The allowlist is a second, independent layer, not the primary safety mechanism — the primary mechanism is architectural (point 2 below): sanitized text is never interpolated into a filter/argv string at all, so no character reaching this far could ever terminate a drawtext option or invoke an expansion regardless of what the allowlist admits. Given that, the allowlist is kept narrow anyway, as ordinary defense-in-depth: colon (:) and percent (%) are excluded (drawtext's own option separator / expansion syntax), and so are',,, and;(not filter-syntax-significant inside a quoted value, but not essential to a lower-third/title/callout either, so excluding them keeps the surface small). Emoji are also deliberately excluded (outside\p{L}/\p{N}/\p{M}): the configured overlay font is not guaranteed to carry emoji glyphs, so admitting them risks silent tofu-box rendering. Combining marks (\p{M}) ARE admitted so NFC-normalized text in scripts without precomposed forms (e.g. Devanagari vowel signs) survives sanitization instead of being silently mangled character-by-character.- Text never appears in the ffmpeg argv or filter string at all. Each overlay's sanitized
text/subtext is written to its own scratch file (
{scratch}/ov-{slot}.txt—DrawtextFilterBuilder.MainTextSlot/SubtextSlotare the single source of truth both the writer and the filter-string builder use for slot numbering) and referenced via drawtext'stextfile=option, never an inlinetext=value — so no drawtext metacharacter (:,',\,%) in the text can ever terminate or inject into the filter string, because the text literally never appears in that string. expansion=noneon every drawtext filter — disables drawtext's own%{...}expansion syntax (which can readpts/localtime/metadataor run%{eif:...}expressions) as defense-in-depth, even though the text is already sanitized and file-based.- Every geometry/timing number is computed in C# (
DrawtextFilterBuilder, from the probed frame size and the already-resolved output-timeline window) and formatted viaFfmpegArgvFormat.Number— culture-invariant, exactly likeEncodeReencodeAsync's existingBetweenTerms()— before being interpolated into the filter string. The model never supplies any of this directly; its only contributions are an offered placement id and a handful of already-resolved enum words.
DrawtextFilterBuilder.BuildFilterChain renames the cut stage's output label from [vout] to
[vcut] only when there is at least one overlay to draw (when EnableGraphics=false, or
EnableGraphics=true but zero overlays survived resolution, the label plumbing is untouched — the
cut stage outputs directly to [vout] exactly as before Phase 3), then chains one drawbox
(semi-transparent background, enable='between(t,start,end)', skipped entirely when
OverlayBoxColor = "none") plus one drawtext per overlay (plus a second smaller drawtext when
Subtext is non-empty), with the LAST overlay's final filter becoming the new [vout] that
-map continues to reference. FilterComplexScriptThreshold's condition now also checks the
built filter STRING LENGTH (spans.Count > 64 || filterComplex.Length > 4000), since overlays can
make one long filter string even with very few cut segments.
Soft-failure discipline: the cut must never become hostage to graphics
Every graphics-specific failure mode degrades to "no graphics applied", never to a failed compile — the cut is the primary deliverable:
| Situation | Outcome |
|---|---|
GraphicsPlan configured but unresolvable/invalid JSON | Cut proceeds with no graphics; graphics.reason records why |
GraphicsPlan not configured at all (null) | Cut proceeds with no graphics; no error |
An overlay's PlacementId not in OfferedPlacementIds | That ONE overlay dropped (unknown_placement_id); the rest still apply |
More overlays than MaxOverlays | Excess dropped (max_overlays_exceeded), in original order |
| Sanitized text ends up empty | That overlay dropped (empty_text_after_sanitization) |
| Placement's window falls entirely in a cut gap | That overlay dropped (cut_away) |
drawtext filter unavailable in this ffmpeg build | ALL overlays skipped; graphics.unavailable = true |
EnableGraphics=true with Mode=StreamCopy | Step FAILS GRAPHICS_REQUIRE_REENCODE — the one graphics failure that IS a hard failure, since it is a pure config error caught before any resolution work, not a soft runtime condition |
The one exception above (GRAPHICS_REQUIRE_REENCODE) is deliberate: drawtext/drawbox filters have
no stream-copy equivalent, so this is a workflow-author config mistake to fix, not a runtime
condition to degrade around.
VideoCompileStepExecutor.BuildEdl/the step's output summary JSON gain a graphics block —
present ONLY when EnableGraphics=true (when false, the EDL/output shape is byte-identical to
the pre-Phase-3 compile path):
{
"graphics": {
"enabled": true,
"applied": true,
"appliedOverlayCount": 1,
"droppedOverlays": [{ "placementId": "p7", "reason": "unknown_placement_id" }],
"unavailable": false
}
}
drawtext availability probe
drawtext needs libfreetype (Alpine's font-dejavu package, which the WorkflowEngine Dockerfile
now installs alongside ffmpeg) and a font file
(VideoEditingOptions.FontFilePath, default /usr/share/fonts/dejavu/DejaVuSans.ttf, wired
through VIDEO_FONT_FILE). A one-off runtime check (ffmpeg -hide_banner -filters, checking for
drawtext in the output) is cached for the process lifetime — never re-probed per step. If
unavailable (an unrebuilt image, or a dev environment missing the font package), ALL overlays are
skipped and graphics.unavailable = true is recorded, but the cut video is still produced
successfully — never fails an otherwise-successful encode over a missing font/filter.
Not built by Phase 3
Text/box drawtext-drawbox overlays with fade in/out, OR a rendered Remotion asset overlay (see
Rendered-asset overlays above) — Phase 3 does NOT apply Ken-Burns
zoompan (Phase 1's KenBurnsCandidate remains identification-only), does NOT burn in subtitles, and
does NOT do transitions between cuts. See Explicitly not built below, which
is unchanged by this phase except for graphics moving out of "not built" and into this section.
Background music
An optional background-music bed, mixed under the dialogue during VideoCompile's encode and
ducked automatically during non-speech windows, planned by a fourth built-in agent that picks among
uploaded tracks — the same structural shape Motion graphics (Phase 3)
established: a deterministic candidate list from VideoAnalyze, an agent that chooses among opaque
offered ids plus a handful of enum words, and VideoCompileStepExecutor alone resolving those words
to real ffmpeg behavior. Off by default (VideoAnalyzeStepConfig.OfferMusicTracks = false,
VideoCompileStepConfig.EnableMusic = false) — EnableMusic = false leaves the compile path
byte-identical to the pre-music behavior.
Why a deterministic volume envelope, not sidechaincompress
MusicMixPlanner/MusicMixFilterBuilder duck the music bed via a deterministic, keyframed
volume=eval=frame envelope computed from the analysis artifact's own silence gaps/transcript
segments — never a runtime audio-level compressor (sidechaincompress). Two reasons this codebase
deliberately does not use a sidechain compressor here:
- The analysis artifact already carries silence gaps (available with zero dependency on ASR) and
transcript segments — exactly the physically-grounded "speech has actually stopped" signal
VideoCompileStepExecutor.ExtendSegmentEndTowardNextSilencealready trusts over ASR boundaries — so there is no need to infer ducking windows from the waveform at encode time at all. - A sidechain compressor's behavior depends on the actual waveform at encode time, so nothing could
assert an exact filter string for it, explain "the music was lifted in these 3 windows" in the
step's own output JSON, or guarantee sane behavior on a source whose dialogue track already has
music baked in. The deterministic envelope, by contrast, is something
MusicMixPlanner.PlanLiftWindowsdecides entirely in C# andMusicMixFilterBuilder.BuildVolumeExpressionturns into an EXACT, assertable ffmpeg filter string.
Candidate discovery (VideoAnalyze)
When OfferMusicTracks = true, VideoAnalyzeStepExecutor enumerates every audio/* project file as
an m{n} music-track candidate (view.musicTracks), capped by MaxMusicTracks (default 20). This
is project-level, not per-source — unlike every other candidate list this feature offers, one
candidate list regardless of how many source clips the step analyzed. Candidates are never
ffprobed here — the fit/duration policy is resolved server-side at compile time regardless of a
candidate's exact length, so probing every candidate here would only cost N downloads for a list the
agent picks at most one item from. A ListFilesAsync failure degrades to zero candidates; this
never fails the step. view.musicTracks is shown as ONE ATOMIC ARRAY, deliberately not gated on
VisualDetail (music has nothing to do with visual detail) — it is dropped as a whole array, after
detail has already degraded all the way to None, before the per-item drop loop ever runs, mirroring
the discipline Motion graphics (Phase 3) established for view.placements.
"Offered" means exactly "the whole musicTracks array survived to the final view" —
VideoAnalysisArtifact.OfferedMusicIds is empty whenever it was suppressed for budget.
m{n} id isolation
Music-track ids (m{n}) are their own namespace, separate from shot/silence/segment ids
(s{n}/g{n}/t{n}), placement ids (p{n}), and look-group ids (k{n}). OfferedMusicIds is its
own separate list — a music-track id must never be validated against OfferedIds/
OfferedPlacementIds and vice versa — and, like placement/look-group ids, is deliberately NOT
resolvable by VideoCompileStepExecutor.BuildIdTimeIndex: a Keep span naming an m{n} id fails
UNKNOWN_ID exactly like any other id that index does not contain.
AgentType.MusicSupervisor and MusicPlanOutput
An ordinary StepType.Agent step, following the exact precedent VideoStoryEditor/
MotionGraphicsPlanner set. Given the story editor's decision (or the same bounded view) plus
view.musicTracks, it picks AT MOST ONE track plus a few coarse settings:
public class MusicPlanOutput
{
public string TrackId { get; set; } = ""; // must be in OfferedMusicIds; empty = no track suits the edit
public string Intensity { get; set; } = ""; // Quiet | Balanced | Feature — never a dB number
public string Ducking { get; set; } = ""; // Off | Light | Normal | Heavy — never a dB number
public string Fit { get; set; } = ""; // LoopToFit | PlayOnce
public string Reason { get; set; } = "";
public string PlanRationale { get; set; } = "";
}
The same rushcut invariant extended again: every property is a plain string, so there is no
numeric/time-bearing CLR type to even ban — enforced by MusicPlanOutputInvariantTests. The model's
only contributions are an opaque TrackId drawn from the set it was actually offered plus the three
enum-word choices; VideoCompileStepExecutor alone resolves those to dB levels/ffmpeg behavior. This
agent is entirely optional: VideoCompileStepConfig.MusicTrackProjectFileId, set directly by the
workflow author, delivers the whole capability (a fixed track, default settings) without this agent
at all. Same minimal read-only project-context + FailWorkflow tool scope as VideoStoryEditor — no
render/sandbox escape hatch the way MotionGraphicsPlanner has, since there is no media for this
agent to produce itself.
Resolution: ResolveMusicAsync, soft-failure throughout
VideoCompileStepExecutor.ResolveMusicAsync runs only when EnableMusic = true, and — like every
other stage in this feature — never fails the compile, only degrades to "no music applied":
| Situation | Outcome |
|---|---|
MusicPlan configured but unresolvable, or not valid JSON | Dropped (plan_unresolved / plan_invalid_json); falls through to MusicTrackProjectFileId if set, else no music |
Plan's TrackId not in OfferedMusicIds, or names no known candidate | Dropped (unknown_track_id); same fallthrough |
Resolved project file not found in the project, or not an audio/* mime type | Dropped (track_not_in_project / track_not_audio); no music |
| Track download or ffprobe fails, or reports zero duration/no audio stream | Dropped (track_download_or_probe_failed); no music |
This ffmpeg build's amix filter has no normalize option | music.unavailable = true; no music (the cut still succeeds) |
The amix normalize-option probe (IsAmixNormalizeAvailableAsync) is cached for the process
lifetime, mirroring IsDrawtextAvailableAsync's pattern exactly — normalize=0 is load-bearing, not
cosmetic: without it amix silently halves every input's level, including the dialogue track, so a
missing option must degrade ALL music rather than risk quietly reducing dialogue loudness.
Resolution order for which track plays: MusicPlan (the agent's choice) first, then
MusicTrackProjectFileId (the deterministic workflow-author-configured fallback) if the plan did not
resolve to a usable track, else no music at all.
Enum words to ffmpeg behavior
Intensity/Ducking/Fit are resolved entirely server-side, exactly mirroring how
Motion graphics resolves Duration/Emphasis:
Intensity(Quiet/Balanced/Feature) → bed level in dBFS viaMusicBedQuietDb/MusicBedBalancedDb/MusicBedFeatureDb(defaults-26/-20/-14), clamped[-40, -6].Ducking(Off/Light/Normal/Heavy) → attenuation below the bed while dialogue is present, viaMusicDuckLightDb/MusicDuckNormalDb/MusicDuckHeavyDb(defaults-6/-11/-18), clamped[-30, 0].Off(from either the model orVideoCompileStepConfig.MusicDucking = MusicDuckingMode.Off) collapses lift-window planning entirely — a flat ducked bed throughout, no trapezoid expression.Fit(LoopToFit/PlayOnce) →LoopToFitadds-stream_loop -1to the music input and trims to the edit's exact frame-quantized length;PlayOncetrims to the track's own length when shorter than the edit, no loop.
MusicMixPlanner: where to lift the bed
MusicMixPlanner.PlanLiftWindows (pure, no I/O, no ffmpeg — the OverlayPlacementBuilder precedent
for this feature's other deterministic planning code) decides WHERE, on the compiled edit's own
OUTPUT timeline, the bed should rise back toward its unducked level. Basis selection, in order — the
first one with any data wins:
silenceGaps(preferred) — the artifact's own detected silence spans, mapped through the sameMapSourceWindowToOutputhelper Motion graphics uses for placement windows.speechComplement— else, the gaps BETWEEN transcript segments, computed per source clip.noSpeechDetected— else, a single lift window spanning the WHOLE output: there is no dialogue anywhere in the kept edit, so the bed should not be needlessly ducked for the whole video.
Candidate windows are merged (closer together than MusicLiftMergeMs), dropped below MinMusicLiftWindowMs
(and always below twice the gain ramp, so a lift too short to fully ramp never reads as pumping), and
capped at MaxMusicLiftWindows (longest kept, re-sorted chronologically).
MusicMixFilterBuilder.BuildVolumeExpression turns the plan into an exact volume=eval=frame
expression: zero lift windows collapse to the bare ducked-gain constant; one window is a single
trapezoid (0 outside [start, end], ramping linearly to 1 across MusicDuckRampMs INSIDE each end of
the window, so a lift is never above the ducked level exactly at a boundary); more than one window
nests binary max(...) calls (ffmpeg's eval has no n-ary max). BuildMixStage's final amix uses
normalize=0 (see above), duration=first (pins the mixed output's length to the DIALOGUE input, so
a looped/infinite music input can never extend the file), and dropout_transition=0 (avoids a gain
re-ramp when the music branch ends before the dialogue does, for PlayOnce with a track shorter than
the edit).
Skipped entirely when the output has no dialogue audio
When the compiled output has no dialogue audio at all — a single audio-less source clip, or (in a
multi-source compile) a mix where NOT ONE referenced clip has an audio
stream — there is nothing to duck against. VideoCompileStepExecutor computes
hasDialogueAudioInOutput (isMultiSource ? anySourceHasAudio : sourceHasAudio) once, after source
download/probe, and threads it into ResolveMusicAsync as hasDialogueAudio. That flag folds into
the SAME duckingOff switch that already collapses lift-window planning to a flat, undocked bed
level (duckingOff = configDuckingOff || !hasDialogueAudio) — so silence-gap/speech-complement
ducking windows are never planned against dialogue that will not exist in the output. music.duckBasis
records "no_dialogue_audio" explicitly in this case (distinct from "none", which means ducking was
simply turned off by config/the model while dialogue audio does exist), and
dialogueHeadroom reports {"applicable": false, "reason": "no_dialogue_audio_in_output"} instead of computing a headroom number against Phase 1 loudness data
for audio that was dropped from the output entirely. The music bed itself is unaffected by this — a
track can still be mixed in as the entire soundtrack of an otherwise-silent edit; only the
speech-aware ducking behavior is skipped.
Narration is the exception. When the output has no dialogue audio but does carry voiceover lines,
those lines ARE the speech to duck under: duckingOff stays false (unless config or the model turned
ducking off), the bed — the LIFTED level — sits NarrationGapLiftDb (12 dB) above the intensity's bed
(capped at -6 dB), and PlanLiftWindows runs with no silences or segments (noSpeechDetected, the
whole program lifted) minus the voiceover windows, so the music plays full between lines and ducks by
the ducking word under each one. music.duckBasis is "narration" and bedBasis is
"narration_gaps". Balanced/Normal: -8 dB between lines, -19 dB under them, against the constant
-20 dB the bed held before.
Review evidence: dialogueHeadroom
VideoCompileStepExecutor.BuildDialogueHeadroom computes deterministic evidence for
AgentType.VideoReviewAgent's StepType.ReviewLoop step: the duration-weighted mean dialogue RMS
across the KEPT spans only (from Phase 1's per-shot VideoAnalysisShotAudio.RmsDbfs) against the
resolved ducked-music level — a hard, server-computed headroom number, not something a model
estimates from audio it cannot hear:
{ "applicable": true, "meanDialogueRmsDbfs": -22.4, "duckedMusicDbfs": -31.0, "headroomDb": 8.6 }
applicable: false when no kept shot carries a Phase 1 audio descriptor (AnalyzeAudioLevels was
off/degraded) or, per the previous section, when the output has no dialogue audio at all.
EDL / output JSON shape
Present only when EnableMusic = true (byte-identical to the pre-music compile path otherwise):
{
"music": {
"enabled": true, "applied": true, "unavailable": false,
"source": "plan", "trackId": "m1", "trackName": "ambient-bed.mp3",
"intensity": "Balanced", "ducking": "Normal", "fit": "LoopToFit",
"bedDbfs": -20, "duckedDbfs": -31,
"trackDurationSec": 42.0, "outputDurationSec": 96.3, "loops": 3,
"playEndSec": 96.3, "fadeInSec": 1.5, "fadeOutSec": 2.5,
"duckBasis": "silenceGaps", "liftWindows": 4, "liftCoveragePct": 18.2, "speechCoveragePct": 71.4,
"dialogueHeadroom": { "applicable": true, "meanDialogueRmsDbfs": -22.4, "duckedMusicDbfs": -31.0, "headroomDb": 8.6 },
"dropped": []
}
}
source is "none" / "plan" / "config" (which of MusicPlan/MusicTrackProjectFileId
actually supplied the track); dropped is a list of {reason, trackId} entries recording every
soft-failure the table above allows, empty when music applied cleanly.
Cutting to the beat
A music-driven reel wants its cuts on the beat and its strongest shot on the drop. Three pieces make that possible without the editor ever naming a time.
Measuring the music (MusicBeatAnalyzer)
Pure C#, no model, no new dependency (Services/Video/MusicBeatAnalyzer.cs). The track is decoded
once through the existing IAudioExtractor (16 kHz mono 16-bit WAV, read by
WavRmsSampler.ReadMonoSamples), then:
- Onset strength — the positive frame-to-frame flux of log energy (10 ms hop, 40 ms window, floored 60 dB under the loudest frame), with its ±0.25 s local mean removed. Each frame's flux is timed at its window's END, where the new energy entered.
- Tempo — the autocorrelation peak of the onset envelope between 70 and 180 BPM (parabolic
interpolation between lags), then refined together with the beat phase by maximizing the
mean onset strength sampled on the beat grid (±1.5 BPM in 0.01 BPM steps, one-hop phase steps).
Confidenceis how far that grid's onset mean stands above the envelope's own mean. - Energy — the RMS level of every beat (each window opens a quarter beat early, so an attack a few milliseconds before the grid line belongs to its own beat). The drop is the beat whose following ~4 s outweighs the preceding ~4 s the most — the earliest such jump within 80% of the biggest, and at least 3 dB. The outro is the beat after the drop where everything that follows sits furthest (at least 4 dB) under the preceding ~4 s.
- Bars are four beats, aligned so the drop is a downbeat (with no drop, on the beat offset carrying the most onset strength).
Golden fixture: "The Last Point" (project file 226a445b-9957-4552-8c99-4736746bbd0b in the local
stack). The numpy prototype it was cut with found 134.25 BPM, phase 0.115 s, drop 13.9 s, outro
~42.6 s over the reel's first 49 s. This analyzer measures 135.0 BPM (identically on the whole
124.6 s track and on 30/49/60 s prefixes), drop 13.93 s and, over the first 49 s, outro 42.8 s; it
recovers a synthetic 134.25 BPM click track to within 0.3 BPM, so the 0.75 BPM gap is attributed to
the prototype's coarser whole-lag envelope. MusicBeatAnalyzerTests holds synthetic click tracks
(always run) and the real track (gated on REELBOLT_BEAT_FIXTURE).
Showing the editor (AnalyzeMusicBeats)
VideoAnalyzeStepConfig.AnalyzeMusicBeats (default false, byte-identical) measures every offered
music track up to MaxBeatAnalyzedTracks (6; each costs one download and decode) and the track named
by MusicBeatTrackProjectFileId — the track the edit is cut to, usually the compile's
MusicTrackProjectFileId. It is a field of THIS step rather than read from the compile step's
config because the step cache keys a step on its own config only. Facts land in the artifact
(VideoAnalysisMusicCandidate.Beat, VideoAnalysisArtifact.Music) and in the view as words and
seconds (MusicBeatProbe.ViewNode):
"music": { "name": "the-last-point.mp3",
"beat": { "bpm": 135.0, "tempo": "fast", "beatSec": 0.444, "barSec": 1.778, "lengthSec": 124.6,
"grid": "steady", "drop": "bar 8, at 13.9 s", "dropSec": 13.93,
"outro": "bar 24, at 42.8 s", "outroSec": 42.83 } }
The VideoStoryEditor prompt (both copies) gained a separate section, "Cutting to the beat (only
when the brief asks)": plan runs in whole bars, put the strongest moment right after the drop by
letting the runs before it add up to its time, end around the outro — and still choose only ids:
VideoEditDecisionOutput and its no-timestamp invariant are untouched.
Snapping the cut (BeatSync)
VideoCompileStepConfig.BeatSync (VideoBeatSync: Off default and byte-identical, Beat, Bar)
runs right after the cut list is final (padding, coalescing, caps, seam bridges) and before the
output timeline is built, so everything downstream — music, narration, overlays, inserts, the EDL —
sees the snapped program. It needs EnableMusic; the track is the music plan's offered choice, else
MusicTrackProjectFileId, and its grid comes from the analyze step's measurement when there is one
(gridSource: "analysis"), else it is decoded and measured here ("measured"). Background music
starts at program time 0 and loops to fit, so the grid is the track's own, repeated per loop
(BeatSyncPlanner.LoopedGrid).
BeatSyncPlanner.Plan walks the spans in program order. Each span's cut moves EARLIER, to the last
grid line (bars, or beats) that still leaves the span at least one beat long — only the TAIL of a
span is trimmed, never extended past its own footage. Bar falls back to a beat when no bar line
fits; with neither, the span is kept whole and the next cut re-aligns. When the drop is at least two
bars in and falls inside a span at least a beat after its start, that span is cut exactly on the
drop, so the next shot lands with it. Contiguous kept shots were already merged into one span, so
only real cuts move. Frame quantization afterwards moves a cut by at most one frame; an overlapping
seam transition (TransitionPolicy) shifts later cuts by its overlap, so pair beat sync with hard
cuts for exact alignment.
Narration is never cut short by a snap. With EnableVoiceover, the planner first reads the
Voiceover step's own output (VideoCompileStepExecutor.NarrationSpanFloors: each ok line's
anchorId and measured durationSec) and gives every span a floor — from the span's start to the
end of the last narration line anchored inside it (a line starts at its anchor or right after the
previous line in that span, whichever is later), plus 0.25 s. A snapped cut must land at or after
that floor; when no grid line fits before the span's own end, the span is kept whole. Found live: a
narrated reel cut on the bar trimmed 3.75 s of tails and 5 of its 10 lines ran onto the next shot
(voiceover.fit.overflowLineCount 5, worst 1.47 s). beatSync.narrationProtectedSpans counts the
spans that carried a floor.
The output summary and the EDL carry, only when BeatSync is not Off:
"beatSync": { "mode": "Bar", "applied": true, "track": "the-last-point.mp3", "gridSource": "analysis",
"bpm": 135.0, "beatSec": 0.4445, "barSec": 1.7779, "dropSec": 13.934, "dropAligned": true,
"keptWholeSpans": 0, "trimmedSec": 1.42,
"cuts": [ { "segment": 0, "outputEndSec": 3.2, "snappedEndSec": 3.044, "trimmedSec": 0.156, "grid": "bar" } ] }
or applied: false with a reason (music_not_enabled, no_music_track, track_not_in_project,
track_not_audio, no_beat_grid, beat_sync_failed) and the cut exactly as the editor made it.
Seam transitions and the program envelope
VideoCompileStepExecutor can apply a short transition treatment at each cut-seam between two kept
spans, and a fade at the very start/end of the whole compiled program — both entirely
deterministic, selected from the same measured shot/seam evidence surfaced to VideoReviewAgent
(see Review evidence below), never from an agent. No step in this pipeline can
request, add, remove, lengthen, or shorten a transition: VideoStoryEditorAgent's prompt states
outright it "cannot create, request, or describe a transition, fade, dissolve, or effect of any
kind," and VideoReviewAgentImpl's prompt is told the same choice is made "deterministically by the
compile step" and to never phrase a fix as "add a fade" or "soften that cut" — only ever as a
different choice of WHICH SPANS to keep. This is the same "rule table, not a model" discipline
Motion graphics's Duration/Emphasis words and Background
music's Intensity/Ducking/Fit words already established for this feature —
except here there is no agent step in the loop at all, deterministic end to end.
Inter-cut transitions vs. the program envelope
Two independent things, controlled by separate config, and easy to conflate:
- Inter-cut transitions happen at every internal seam between two kept spans — a hard cut, a
brief audio-only declick ramp (
AudioSeamRampMs), a short crossfade dissolve (DissolveMs), or a dip-to-black/dip-cut (DipToBlackMs/DipCutMs) — chosen per seam from that seam's own measured properties, capped byMaxTransitionMs/MaxTransitionRatioPctso a transition can never eat a meaningful fraction of either neighbouring segment. - The program envelope is the single fade-in at the very start and fade-out at the very end of
the WHOLE compiled file — video (
ProgramFadeInMs/ProgramFadeOutMs) and audio (ProgramAudioFadeInMs/ProgramAudioFadeOutMs) tracked separately, since a video dip-to-black and an audio fade need not move in lockstep. This has nothing to do with any internal seam; it exists purely so a finished piece never starts or ends on a hard, un-eased frame — the same weightVideoStoryEditorAgent's own "The opening and the closing" prompt section already places on the first and last kept span, applied here at the encode instead of the edit-decision stage. The audio envelope never applies to narration: voiceover lines are mixed after it, and a last line running past the picture holds the final frame until it has finished — seevoiceover.md"Narration and the end of the program".
TransitionPolicy
VideoCompileStepConfig.TransitionPolicy (a VideoTransitionPolicy enum) gates the whole inter-cut
system: at its most permissive setting — "Auto", what all three video-derush-edit* templates now
request — the rule table is free to apply whichever treatment a seam's own measurements call for;
turned off entirely, every seam falls back to a hard cut and the compile path is byte-identical to
the pre-transition behavior, the same "off by default, byte-identical without it" guarantee
EnableGraphics/EnableMusic already give this feature. The full set of intermediate levels the
enum exposes belongs to the sibling change that introduces VideoTransitionPolicy itself; the
contract fixed here, and depended on by every other piece of this feature (the prompts above
included), is that every policy level except Editor selects a treatment FROM measured data, and
Editor lets the editor choose only by WORD — see Editor-chosen transitions.
Editor-chosen transitions
TransitionPolicy = Editor (appended last) hands each seam to the editor agent. Every
VideoEditKeepSpan may carry Transition — a preset WORD — and TransitionSpeed — Snap, Quick,
Smooth or Slow — naming how that span comes IN; both are strings, so the rushcut invariant (no
number in the decision) holds, and PickupPlanOutputInvariantTests pins the widened property set.
EditorTransitionCatalog alone resolves them: 56 presets map one-to-one onto ffmpeg's built-in
xfade transitions (Fade, FlashWhite, Zoom, Pixelize, SlideLeft, PushUp, RevealRight,
WipeDown, SmoothLeft, CircleOpen, BarnDoorOpen, SliceLeft, WindRight,
SqueezeHorizontal, …; a raw xfade name is accepted too), DipToBlack takes the DipToBlack
treatment and every other preset the Dissolve one with its own xfade name, Cut is a hard cut,
and the speeds are 0.2 / 0.35 / 0.6 / 1.0 s (Quick when a preset names none).
A resolved span takes the choice of the kept span it starts with (ResolvedSpan.OriginKeepFromIds[0]);
a seam whose span names nothing, or names an unknown preset (reported in
transitions.editor.unknownPresets), takes the Auto rule table. An editor choice is NOT exempt from
the safety passes: no xfade demotes it to DipCut, a seam touching a chroma-plate span is still
demoted, and the per-seam (half the shorter span, MaxTransitionMs) and neighbour-sum (60%) overlap
clamps still apply. It IS exempt from the density cap (MaxTransitionRatioPct), which only downgrades
rule-table seams: an editor that put a transition there meant it. Its seams carry the rule
E:<Preset>:<Speed>, and transitions.editor reports chosenCount and every chosen seam
(seam, choice, treatment, transition, durationSec). The VideoStoryEditor prompt (both
copies) lists every preset and teaches restraint — most seams are cuts, directional moves keep one
direction, hard hitters go on a beat. Under any other policy the two fields are ignored.
SectionBreakGapMs and silence-anchored seams
A seam that falls on a long-enough silence gap (SectionBreakGapMs) reads as a genuine section
break rather than an ordinary mid-sentence cut, and the rule table treats it differently from a seam
with no silence on either side — the same kind of distinction seamCheck
already exposes to the reviewer via cutOutMotion/cutInMotion for motion, applied here to silence
instead.
MaxTransitionSegments: the filtergraph-buffering guardrail
A segmented single-source encode already writes one select/aselect filtergraph entry per kept
segment (see MaxSegments above); a crossfade-style transition needs to buffer and blend TWO
adjacent segments at once instead of switching between them instantaneously, which multiplies the
filtergraph's memory/CPU cost per transition rather than per segment. MaxTransitionSegments caps
how many segments a compile is willing to apply transitions across at all — above the cap,
transitions are skipped for the WHOLE compile (every seam falls back to a hard cut) rather than
risking an ffmpeg process that OOMs or times out on a long, heavily-cut edit. This mirrors
MaxSegments's own "switch to a scratch-file filtergraph above ~64 segments" guardrail in spirit: a
structural limit protecting the ffmpeg process, not a quality knob.
Review evidence
VideoCompileStepExecutor's output JSON surfaces transitions (policy, appliedCount, and a
treatments breakdown) and programFade alongside the existing sentenceCheck/graphics/music
blocks, plus three more deterministic checks that arrived alongside the transition system:
openingCheck (the sentenceCheck mirror for the FIRST kept span instead of the last),
seamCheck (every inter-cut seam's measured look/motion continuity, both an aggregate count and an
itemized list of the notable ones), and pacing (segmentCount/meanSegmentSec/
medianSegmentSec/shortSegmentPct). All of these are computed once per compile and read — never
re-derived — by AgentType.VideoReviewAgent's ReviewLoop step exactly like sentenceCheck/
graphics/music already are; see VideoReviewAgentImpl's prompt for the exact score caps and
remediation rules each one drives. When the transitions node is absent entirely, transitions were
switched off for that workflow and the reviewer is told explicitly not to penalize hard cuts.
Semantic visual dimensions (Phase 4)
Seven deterministic dimensions (D1-D7) added on top of Phase 1's scene/visual analysis, plus
prompt/contact-sheet/persistence changes to Phase 2's vision captioning. D1-D4 and D6 are free
— derived from the same low-res grid data Phase 1 already samples, so they default on. D5
(audio character) rides AnalyzeAudioLevels the same way. D7 (backlit) is a byproduct of D6's
region data, also free. Only sharpness (a D-adjacent dimension, not part of D1-D7's own numbering)
costs a genuinely new ffmpeg invocation per measured shot, so it alone stays opt-in.
Every new stage below follows Phase 1's original discipline: degrade, never fail the step. A classification that cannot be computed (near-monochrome frame, too few audio windows, no clean letterbox bars) simply omits that field or id — it never throws out of the analysis pass, and it never blocks silence/shot detection, transcription, or any other deterministic stage from completing.
| # | Dimension | Formula (informal) | Gate | Documented limitation |
|---|---|---|---|---|
| D1 | Color temperature (Warmth/Tint/ColorTemperatureClass) | Warmth = clamp((meanR − meanB) / 0.25, −1, 1); Tint is the same shape against G vs. (R+B)/2. ColorTemperatureClass is Warm/Cool at |Warmth| ≥ 0.20, else Neutral; a near-monochrome frame (SaturationMean < 0.05) is always Neutral regardless of Warmth | AnalyzeColorGrading | A single dominant colored object (not the lighting) can skew the whole-frame mean; this is a frame-average heuristic, not a white-balance measurement |
| D2 | Tone curve (BlackPoint/WhitePoint/ToneClass) | 5th/95th percentile of the luma histogram (nearest-rank). ToneClass order is load-bearing — an actual exposure defect always outranks a stylistic read: Blown (clipped highlight ratio > 5%) → Crushed (crushed black ratio > 5%) → Flat (dynamic range < 0.45 and black point > 0.10 — lifted blacks + compressed range, i.e. log/ungraded) → Contrasty (dynamic range > 0.75 and black point < 0.06) → Normal | AnalyzeColorGrading | Flat on a whole look group is a property of the SOURCE footage (ungraded log), not a per-shot defect — both agent prompts say so explicitly |
| D3 | Saturation character (SaturationClass) | Pure threshold projection of the pre-existing SaturationMean: Muted (< 0.18) / Natural / Vivid (> 0.42) | AnalyzeColorGrading | Same mean-based coarseness as D1 |
| D4 | Look grouping (view.lookGroups, ids k{n}) | Six-float LookSignature (Warmth, Tint, BrightnessMean, BlackPoint, WhitePoint, SaturationMean) per shot; LookDistance is a weighted L1 distance normalized per-component to 0..1 (weights 0.30/0.10/0.25/0.15/0.10/0.10, summing to 1.0 so Similarity = 1 − Distance lands in 0..1); single-linkage clustering over all pairs (not windowed like near-duplicate grouping, since a shared look deliberately links non-adjacent shots/clips) at LookSimilarityThreshold (default 0.88). LookRank orders group members by ascending distance to the group centroid — rank 0 is the most representative shot | DetectLookGroups | O(shots²) — trivially cheap at the shot counts this feature targets, but would need revisiting at extreme shot counts |
| D5 | Audio character (char under "a") | ShotAudioAnalyzer.Analyze: crest factor (peak − RMS dB), 10th-percentile window RMS as a noise floor, level stability (1 − clamp(stdDev/12dB, 0, 1)), and zero-crossing rate. Classification order is load-bearing (Dialogue is the fallthrough, never a positive claim): Silent (RMS ≤ −50 dBFS) → Music (stable + compressed + low ZCR) → Noisy (high noise floor, low speech ratio) → Ambient (quiet, low speech ratio) → Dialogue | AnalyzeAudioLevels | Deliberately a ZCR/crest/noise-floor heuristic, not a spectral (FFT) classifier — this repo's only audio test fixtures are synthetic sine tones, and thresholds tuned against a 440 Hz sine would pass CI while misclassifying real footage (judgment call 7) |
| D6 | Letterbox/pillarbox (ActiveCrop) | Scans the already-materialized luma frames for rows/columns whose luma stays below a tolerant black threshold (16/255, tolerant of compression noise inside a true matte) across every sampled frame, capped at 40% of the frame dimension (beyond that it reads as a dark scene, not bars) | DetectLetterbox (default true — free, no second luma pass) | Under-reports soft/gradient letterbox edges — the threshold expects a clean black bar, not a feathered one |
| D7 | Backlit candidate (BacklitCandidate) | Byproduct of D6's region data: the center region (R4) is markedly darker than the average of the surrounding border regions (borderLuma − centerLuma > 0.18) while itself being dark (centerLuma < 0.35) | Free whenever regions are computed | Named and surfaced as a candidate, not a claim (mirrors KenBurnsCandidate's precedent) — fed to the vision prompt so a model that can actually see the frame turns it into (or rejects) an actual assessment; Full detail only |
| — | Sharpness (Sharpness, D-adjacent) | Native-resolution, square, centered grayscale patch (FfmpegArgvBuilder.BuildSharpnessPatchArgs, deliberately its own ffmpeg invocation so a fault there can never take Phase 2 keyframe extraction down with it) → discrete 4-neighbour Laplacian → variance over the interior, normalized against a documented heuristic constant. Only ever compared BETWEEN shots of the same source, never as an absolute unit | DetectSharpness (default false — one extra ffmpeg call per measured shot) | Costs real wall-clock/CPU unlike D1-D4/D6/D7, which is why it alone stays opt-in; capped by MaxSharpnessShots, a genuinely step-wide budget (like MaxCaptionedShots) computed AFTER the per-shot analyze loop and BEFORE that source's own duplicate grouping, so a real measured value (when available) — not the neutral placeholder — reaches ComputeTakeQuality's best-take scoring |
k{n} isolation
Look-group ids (k{n}) are a purely descriptive namespace, structurally different from every
other id this feature offers a model. Shot/silence/segment ids (s{n}/g{n}/t{n}) are offered to
VideoStoryEditorAgent and resolved by VideoCompileStepExecutor.BuildIdTimeIndex; placement ids
(p{n}) and music-track ids (m{n}) are separate offered namespaces resolved by their own
executor paths. k{n} is never offered to any agent at all — there is no OfferedLookIds list,
deliberately — and no structured output in this feature (VideoEditDecisionOutput,
MotionGraphicsPlanOutput, MusicPlanOutput) has a field that can name a look group. A Keep span
naming k{n} can therefore only be a hallucination, and BuildIdTimeIndex deliberately excludes
k{n} from its index so such a span fails UNKNOWN_ID exactly like any other id it does not
contain — the same fate as a p{n}/m{n} id named in a Keep span.
View/artifact shape
{
"view": {
"shots": [
{
"id": "s0", "startSec": 0.0, "endSec": 4.2, "durationSec": 4.2, "src": 0,
"v": {
"motion": 12, "move": "Static", "cutIn": "still", "cutOut": "still",
"bright": 58, "colors": ["#3a2c1e", "#c9a876"],
"temp": "Warm", "tone": "Normal", "sat": "Natural", "look": "k0", "crop": [0.0, 0.11, 1.0, 0.78]
},
"a": { "rms": 42, "speech": 71, "char": "Dialogue" },
"c": {
"summary": "A presenter gestures at a whiteboard while explaining a diagram.",
"scale": "Medium", "mood": "Focused", "tags": ["presenter", "whiteboard"],
"style": "flat ungraded log", "issues": []
}
}
],
"lookGroups": [
{ "id": "k0", "shotIds": ["s0", "s1", "s4"], "repShotId": "s0", "cohesion": 91, "temp": "Warm", "tone": "Normal", "sat": "Natural" }
]
},
"meta": {
"look": { "applied": true, "uniform": false, "groupCount": 1 },
"vision": { "mode": "Optional", "applied": true, "captionedShots": 6, "failedShots": 0, "persistedKeyframes": 0 }
}
}
"temp"/"tone"/"sat" and "look" live under a shot's existing "v" node — gated on
AnalyzeColorGrading/DetectLookGroups and VisualDetail exactly like every other Phase 1 field,
same degrade-before-drop discipline. "crop" (D6) appears only when a crop was actually detected —
omitted, not null, when the frame is full-bleed. "char" (D5) is always present on the "a" node
whenever audio levels were analyzed, including for the common Dialogue value — unlike the visual
fields, there is no "absent means nothing to report" reading for D5, since every shot has SOME
audio character. view.lookGroups is capped by MaxViewLookGroups; when the whole analysis is one
uniform look, every per-shot "look" id is suppressed and meta.look.uniform is true instead —
repeating the same group id on every shot would add bytes without adding information.
Progress weighting
VideoAnalyzeProgressPlan's per-source stage list gained SampleSharpness (weight 4, only
counted when DetectSharpness is on) between AnalyzeShots and GroupDuplicates, and the
step-level list gained MatchLooks (weight 2, only counted when DetectLookGroups is on) between
GroupDuplicates/multi-source join and ListMusicCandidates:
| Stage | Weight | Level |
|---|---|---|
SampleFrameGrid | 12 | Per-source |
AnalyzeShots | 10 | Per-source |
SampleSharpness | 4 | Per-source (only when DetectSharpness) |
GroupDuplicates | 2 | Per-source |
MatchLooks | 2 | Step-level (only when DetectLookGroups) |
CaptionShots | 14 | Step-level |
A disabled stage contributes zero weight and is skipped entirely rather than reported as an instant 0%-to-100% jump — the same "only enabled stages count toward the total" rule the original plan established for every other optional stage, so turning a Phase 4 dimension off never distorts the percentages reported for the stages that stayed on.
Vision-phase changes
See Prompt priming, contact sheets, and persisted keyframes (Phase 4) under Vision captioning, and the multi-source hoist fix noted at the top of Executor wiring — both are Phase 4 changes to the Phase 2 captioning path, documented alongside Phase 2 rather than duplicated here.
Tracked screen inserts (Phase 5)
Compositing a Remotion-rendered scene INTO a moving region of the source footage — the canonical
case: a commercial shot with a phone held in frame against a green screen, where an app UI
(rendered as its own Remotion composition) is inserted into the phone's screen area and moves and
warps with the phone as the hand moves, not as a static overlay. Off by default at both ends
(VideoAnalyzeStepConfig.DetectInsertRegions = false, VideoCompileStepConfig.EnableInserts = false); EnableInserts = false leaves the compile path byte-identical to before this phase.
The invariant this phase exists to protect: a motion-tracking transform is inherently
per-frame numeric data — positions, corner coordinates, times. Every one of those numbers is
computed by deterministic C# and consumed by deterministic C#. The model's entire contribution is
(1) an opaque r{n} region id drawn from the set it was actually offered
(VideoAnalysisArtifact.OfferedInsertRegionIds — a separate id namespace, never resolvable by
BuildIdTimeIndex, "offered is stricter than exists" like every other id family) and (2) a
rendered asset it produced itself via a real RenderVideoAndUploadToStorage call, validated
against this execution's own outputFiles prefix — the exact RenderedAssetStorageKey precedent.
ScreenInsert (on MotionGraphicsPlanOutput) has exactly three string properties — RegionId,
RenderedAssetStorageKey, Reason — locked by MotionGraphicsPlanOutputInvariantTests, which
also pins the exact property set so even a new string property is a visible, reviewed decision.
The tracking decision: chroma-plate quads, not general feature tracking
ChromaQuadTracker (WorkflowEngine/Services/Video/ChromaQuadTracker.cs) is a pure, unit-tested
static class. It runs over a dedicated, medium-resolution raw-RGB grid pass (the same
IFrameGridSampler machinery Phase 1 uses, at ~320px on the long axis / 10fps instead of
32x18/2fps) and, per frame: builds a chroma mask by relative channel dominance
(brightness-robust — g*10 > r*13+100 etc., for green/blue/magenta), finds the largest 4-connected
component (BFS), grows it through a second, permissive mask (see "Dual-threshold chroma masking"
below), fits a quadrilateral via the extreme-point method (TL=min(x+y), BR=max(x+y), TR=max(x−y),
BL=min(x−y)), and gates on area ratio, component-vs-quad fill ratio (rejecting L-shapes and
scattered noise) and quad degeneracy. Per-frame quads assemble into tracks (dropout gaps ≤ 3 frames
bridged by linear interpolation; larger gaps or implausible centroid jumps split the track),
corners get a small centered moving-average smooth, and each track carries a confidence plus
qualitative size/motion/aspect descriptors.
The tracking grid preserves the source's aspect ratio
The grid's dimensions are derived from the source's own probed aspect (DeriveGridSize), not
fixed. InsertGridWidth/InsertGridHeight default to 0 = auto; a positive pair is an explicit
override and is honoured verbatim.
This matters twice over. A fixed 320x180 grid squashes a 1080x1920 portrait source vertically
by 10.7x, so (a) vertical corner precision collapses to ±10.7 source px against ±3.4 horizontal —
past InsertOverscan's 0.02 default, which on real footage left a visible green fringe (measured:
0.773% residual green at the fixed grid versus 0.090% at an aspect-matched 202x360), and
(b) every physical distance measured in grid space is wrong by the aspect mismatch. The most
visible casualty was MeanAspectRatio, the one real number offered to MotionGraphicsPlanner:
a phone plate whose true aspect is 0.47 was reported as 1.46 — telling the agent a portrait
screen was landscape.
Auto-derivation fixes the precision half. The aspect half is fixed independently and belongs to
defense in depth: Track takes the source's real sourceWidth/sourceHeight as explicit
parameters used only for physical-distance measurement, so the reported aspect stays correct
even if a workflow author pins a grid that disagrees with the source. 16:9 sources still derive
exactly 320x180, so nothing changes for the common case.
Dual-threshold chroma masking, and what confidence actually measures
The channel-dominance ratio test is brightness-robust; the absolute floor beside it is not.
A single floor of g >= 60 sat in the middle of a real dim laptop panel's own brightness
distribution — that plate's green channel measured 28..88 across the same physical screen — so
the mask captured only 66% of it (8.97% of frame area against a tolerant test's 13.56%), and
the extreme-point fit then fitted a confident, tidy quad to the brighter fragment, leaving a
large triangular strip of raw green exposed in the composite.
Detection is therefore dual-threshold (hysteresis), the standard fix for exactly this:
- the core mask (floor 60, unchanged margins) seeds detection with pixels that are unambiguously the plate;
- the grow mask (floor 24 — where 8-bit channel ratios stop being meaningful at all — with proportionally relaxed margins) extends that seed, and only that seed, by 4-connectivity.
Growth can never jump to a different object: an unrelated green-ish blob is admitted only if it is
physically contiguous with the high-confidence core, and if growth ever did merge scenery the
existing MinFillRatio gate rejects the resulting non-quadrilateral blob exactly as before.
Confidence was coverage × mean fill ratio, with fill clamped to 1.0. That made the second
term a constant in practice (a discretized quad fit normally over-fills slightly, ~1.02–1.07), so
the score was effectively just coverage — and nothing in it could see under-detection, because
every term was computed against what the strict mask happened to find. That is how a visibly wrong
detection reported 0.9996 and sailed past MinInsertConfidence. It is now a product of four
independent axes (the photometric and temporal terms were revised later — see
Self-tuned growth and honest confidence):
| Term | Axis | Catches |
|---|---|---|
coverage | temporal | frames where nothing was detected at all |
meanFillScore | geometric | a fit that does not bound its own component. Two-sided (fill or 1/fill), so a degenerate fit scores low instead of being clamped up to perfect — a 45° plate collapses two extreme points onto the same pixel and fits a triangle at fill ≈ 2.03, which the old clamp rewrote as 1.0 |
meanEdgeContrast | photometric | does each fitted edge sit on a real plate boundary — chroma purity just inside the edge minus just outside it. Replaced meanChromaMargin (|core| / |grown|), which docked a correct dim-plate fit for needing growth and could not tell a fit that leaked into scenery from a good one |
consistency | temporal | the fraction of frames whose four corners agree with their neighbours' temporal median — a physical plate moves continuously, a fit that flickers onto scenery does not |
A quad with two coincident corners is now rejected outright rather than warped. On the real
footage the net effect is: the laptop plate is detected at full extent (13.6% of frame, corners
landing on the screen edge instead of ~100px inside it) and scored 0.63–0.64 instead of
0.9996, while the clean phone plate still scores 0.91–0.92. Note the goal is not to reject
the laptop shot — with the mask fixed its detection is correct — but to make the score mean
something, so raising MinInsertConfidence can actually exclude marginal plates.
Why marker/chroma-based, and what was rejected. General markerless planar tracking (KLT/feature-correspondence homography estimation) was evaluated and rejected for v1, in the same cost-benefit style as the "why ffmpeg is not in the sandbox" decision:
- OpenCV via OpenCvSharp would be the honest way to do markerless tracking, but it is a large native dependency with no musl/Alpine binaries the WorkflowEngine image could consume without building OpenCV from source — hundreds of MB of image growth, a large native attack surface parsing untrusted media (the same class of risk the existing ffmpeg-hardening flags exist for), and a build pipeline burden, for a capability whose robustness could not be validated here against real footage anyway.
- A from-scratch C# feature tracker could not honestly be claimed robust — pyramidal Lucas-Kanade plus RANSAC homography fitting is a genuine CV subsystem, not a helper class.
- A chroma plate is the one target pure pixel statistics detect reliably — and it matches how this shot is actually produced in practice (a phone screen displaying solid green IS the standard on-set practice for exactly this composite). The constraint is stated plainly: the source footage must contain a uniform-color plate (green by default; blue/magenta configurable). Footage without one gets zero offered regions, and the agent is prompted to plan zero inserts.
The tracker's precision is bounded by the tracking grid (~0.3% of each frame dimension at the
auto-derived default — ±6px at 1080p, now equal in both axes because the grid matches the source's
aspect), softened by corner smoothing and by InsertOverscan expanding the insert slightly past
the plate's edge. Raise InsertGridWidth/InsertGridHeight when tighter registration matters more
than the larger grid buffer — but set BOTH, since a half-set pair falls back to auto.
The full numeric track (VideoInsertRegionTrack.Keyframes — per-sampled-frame normalized corner
quads on the source timeline) lives ONLY in the analysis artifact. The bounded view offers
view.insertRegions: {id, shotId, startSec, endSec, conf, size, motion, aspect, color} (+src
when multi-source) — the window is READ-ONLY input exactly like view.placements'
startSec/endSec, conf/size are bucketed words, and aspect is the one number the agent
genuinely needs as input (to render suitably-proportioned content). The array follows the
musicTracks budget discipline (unconditional, atomically suppressible — deliberately NOT the
detail-gated placements discipline, since insert regions come from their own grid pass and are
independent of Phase 1 visual analysis).
Agent surface
No new agent: AgentType.MotionGraphicsPlanner gained an Inserts list on its existing
MotionGraphicsPlanOutput (additive — every pre-existing plan deserializes with zero inserts),
and its prompt (fallback + seeded, verbatim-locked by VideoStoryEditorPromptConsistencyTests)
gained a "Tracked screen inserts" section.
Two things this addition originally missed, both found by running it.
AgentType.MotionGraphicsDirector — the agent that performs the graphics room's structured
synthesis, i.e. the one that actually emits a MotionGraphicsPlanOutput — never got that section,
and its prompt closes by enumerating what to output as "an overlays list ... and a
planRationale". Handed two real tracked plates it rendered both insert scenes, described them in
its planRationale, and emitted an empty inserts list. Separately,
DatabaseSeeder.GenerateMotionGraphicsPlanSchema — the hand-written schema documentation seeded
onto every agent row — was never updated either, so the platform's own schema viewer described a
two-field plan; and because existing rows only had OutputSchemaJson written when the column was
empty, correcting the literal could not have reached an already-seeded deployment. All three are
fixed, and SeededOutputSchemaDriftGuardTests now compares every seeded schema's top-level
property names against its CLR output type so the documentation cannot silently fall behind the
type again. The agent renders the insert content itself through the
same sandbox+Remotion pipeline it already uses for rendered-asset overlays — but OPAQUE (a normal
mp4, no alpha), since the whole rectangular frame is warped to fill the plate.
Compile: the corner-pin recipe
VideoCompileStepExecutor.ResolveInsertsAsync (soft-failure throughout) validates each insert,
maps the track's source window through the cut (OutputTimeline.MapWindowToOutput — clipped to
the FIRST kept portion, like overlays), converts each surviving tracked keyframe to an
output-frame-indexed pixel quad (BuildInsertKeyframes: per-keyframe MapToOutputSec, uniform
downsample to MaxInsertExprKeyframes, centroid overscan expansion), downloads + ffprobe-validates
the asset, and hands ScreenInsertFilterBuilder a fully-resolved numeric description. The
filtergraph per insert — validated end-to-end against a real ffmpeg run during design:
[N:v] scale=(W-2)x(H-2), fps=canonical, pad to WxH with a 1px black border,
perspective sense=destination eval=frame (corner exprs piecewise-linear in `in`),
setpts +outputStart -> warped content
color=white (W-2)x(H-2), pad 1px black border, format=gray,
the SAME perspective exprs, the same setpts -> warped mask
alphamerge(content, mask) ; overlay at 0:0, enable='between(t,start,end)'
The 1px border is load-bearing: perspective edge-clamps out-of-range source coordinates, so an
unbordered warp smears content across the whole frame outside the quad; bordering both the content
and an all-white mask makes everything outside the warped quad black, which alphamerge turns
into transparency (perspective itself supports no alpha format — that is why the mask branch
exists). The insert stage renders BEFORE text/asset overlays (screen content is scene content;
lower-thirds paint on top), its inputs sit between the asset-overlay inputs and the music input
(preserving both existing index mappings), and corner expressions are piecewise-linear
if(lt(in,f),a+(in-f0)*s,...) chains over the filter's per-frame in variable — every literal
through FfmpegArgvFormat.Number. Not one model-originated character reaches the filter string.
in is 1-based (ffmpeg evaluates it as the frame count plus one), so keyframe k is written
at in = k + 1; until that was known every insert ran one frame ahead of its plate — see
Tracking quality on the customer clips.
Soft-failure table
| Situation | Outcome |
|---|---|
EnableInserts=true with Mode=StreamCopy | Step FAILS INSERTS_REQUIRE_REENCODE — pure config error, mirrors GRAPHICS_REQUIRE_REENCODE |
No GraphicsPlan configured / plan unresolvable / invalid JSON | No inserts applied; inserts.reason records why |
RegionId not in OfferedInsertRegionIds | That insert dropped (unknown_region_id) |
| Same region chosen twice | Second dropped (duplicate_region_id) |
More inserts than MaxInserts | Excess dropped (max_inserts_exceeded) |
Track confidence below MinInsertConfidence | Dropped (confidence_below_threshold) |
Asset key missing or outside this execution's outputFiles prefix | Dropped (invalid_asset_storage_key) |
| Asset download/probe fails | Dropped (asset_download_or_probe_failed) |
| Track window entirely inside a cut gap, or its own source clip contributed no output time at all | Dropped (cut_away) |
perspective/alphamerge missing from the ffmpeg build | ALL inserts skipped; inserts.unavailable = true |
| Window falls entirely inside a crossfade seam's blend region | Dropped (window_consumed_by_transition_overlap) |
The EDL/output summary gains an inserts block (present only when EnableInserts=true):
{enabled, applied, appliedInsertCount, droppedInserts: [{regionId, reason}], unavailable}.
v1 scope, stated plainly
What works: a chroma plate (green/blue/magenta) tracked as a deforming quadrilateral —
translation, scale, rotation, and perspective skew all follow the plate, since all four corners
are tracked independently and the warp re-evaluates per frame. On both encode paths: the
original single-source select path and the segmented concat path used for multi-source
compiles and crossfade transitions.
Inserts on multi-source compiles
Inserts were originally skipped outright whenever the compile routed to the segmented encode, which meant a green-screen clip could not be used in any multi-source edit — the raw flat plate simply appeared in the final cut. That was encode-path plumbing, not a limit of the math: the corner-pin never depended on how the base video was assembled.
OutputTimeline is already source-aware, and both mapping calls pass the tracked region's own
track.SourceIndex — MapWindowToOutput for the on-screen window, MapToOutputSec per corner
keyframe. Both only ever intersect spans belonging to that region's own clip, so an insert
resolves onto exactly the output time that clip contributed. Inertness during other clips'
segments is not a special case: the composite is one overlay gated by
enable='between(t,start,end)' over that window, so during output time cut from a different clip
the overlay contributes nothing. This is the same mechanism OverlayAssetFilterBuilder already
uses to gate an asset overlay to its own on-screen window.
The segmented path's stage order now mirrors the single-source path exactly —
concat → colour grade → screen inserts → text overlays → asset overlays — and insert asset
inputs sit after every source input and every asset-overlay input, before the music and SFX inputs
(all downstream index offsets shifted accordingly).
Geometry needs the letterbox, not just the timeline. A tracked quad's corners are normalized to
its OWN source frame. On the single-source path that is the canvas, so normalized corners times
canvas dimensions is correct. On the segmented path with several clips every span is normalized
with scale=cw:ch:force_original_aspect_ratio=decrease plus a centered pad, so a clip whose
aspect differs from the canvas is letterboxed or pillarboxed and occupies only part of it.
SourceCanvasFit reproduces that same arithmetic per source, and BuildInsertKeyframes maps
through it. Without this the insert composites in the right place in time and the wrong place in
space: caught on a real three-source compile where a 1080x1920 phone clip pillarboxed into the
canvas had its insert drawn roughly three times too wide, spilling across the entire frame. Note
the canvas is the FIRST KEPT span's source, not a fixed 1920x1080 — reorder the edit and the canvas
changes with it, which is exactly why the fit is computed per source rather than assumed.
Overscan still expands about the plate's centroid in the source's own normalized space before the
fit maps it, so InsertOverscan keeps meaning "a fraction of the plate" rather than "a fraction of
the canvas". The mapping is affine, so the order is equivalent.
One case genuinely does not admit an insert: an overlapping seam treatment (dissolve/dip/whip)
blends two spans into the same output frames, and a quad pinned to one clip's tracked geometry
must not paint over a blended frame. OutputTimeline.ClipAwayTransitionOverlaps trims an insert's
window back to its longest clean stretch, and drops it
(window_consumed_by_transition_overlap) only if nothing survives. This is a no-op for every
compile at the default TransitionPolicy = Off.
Deferred, deliberately: markerless tracking of arbitrary regions (needs a real CV dependency — see the decision record above); sub-pixel corner refinement at native resolution; a single tracked region spanning SEVERAL non-adjacent kept spans (the window is still clipped to its first kept portion, exactly as overlays are, so a plate cut into two pieces composites into the first only); an insert painted over a crossfade's blended frames (see above); lighting/color match of the insert to the scene (the content is composited as rendered — no ambient wrap, no relight); motion blur on fast plate movement.
Several of Phase 5's deferrals have since been lifted — non-planar (curved/bent) plates and multi-plate-per-frame tracking, see Non-planar surfaces and multi-plate compositing; occlusion, see Surviving occlusion. The statement below that the composite is a quadrilateral is also no longer true of the default path: it is cut to the plate's actual silhouette, see Silhouette matting. The v1 statement above that "all four corners are tracked independently" so translation, scale, rotation and perspective skew all follow the plate remains exactly true of the single-quad path, which is still what a flat plate uses.
Non-planar surfaces and multi-plate compositing
Two limitations Tracked screen inserts (Phase 5) named as deferred, lifted. Phase 5 fits ONE planar quadrilateral per plate, per frame, which is exactly right for a flat phone screen held in frame and wrong for anything that bends — a curved monitor, a flexed card, a plate wrapped on a cylindrical object. And it takes only the LARGEST chroma component per frame, so two devices held up together resolve to one flickering track instead of two.
Both additions are off or inert by default (MaxInsertPlatesPerFrame = 1,
InsertSurface = Auto with MinInsertCurvature = 0.01 that a flat plate stays well under), and a
plate that measures flat keeps the original single-perspective corner-pin filtergraph.
What a chroma silhouette can and cannot tell you
A uniform-colour plate gives exactly one thing: its outline. That bounds the problem honestly.
- The outline does show a bend. A flat rectangle in perspective has four straight edges; the same rectangle bent about an axis has two curved ones. That curvature is measurable from pixel statistics alone.
- The outline does not show texture foreshortening. Where a point on the interior of a curved surface projects to depends on the surface's 3D pose, and a plate of uniform colour has no landmarks to recover it from. No amount of silhouette processing changes that.
So this is an image-space deformation model, not a 3D reconstruction, and the guarantee it makes is the one the silhouette can actually support: the composited content covers exactly the region the plate covers. That is the error that shows as green fringe, and it is the one that matters. The interior distribution is smooth, temporally stable and plausible, but is not physically exact for a strongly curved or steeply yawed surface. Measured against analytically known cylindrical geometry (a rectangle bent on a circular cylinder, perspective projected):
| Case | Boundary error, flat quad | Boundary error, mesh | Interior error, flat quad | Interior error, mesh |
|---|---|---|---|---|
| Gentle bend, distant camera | 3.9px | 0.01px | 3.9px | 0.8px |
| Strong bend, distant camera | 14.2px | 0.02px | 14.4px | 9.1px |
| Strong bend, close camera | 40.1px | 0.87px | 39.5px | 5.7px |
| Convex bulge (can/bottle) | 19.2px | 0.43px | 24.1px | 19.8px |
| Bend + 25° yaw | 13.1px | 0.51px | 39.6px | 37.4px |
| Bend + 45° yaw | 10.3px | 0.43px | 64.5px | 63.6px |
| Horizontal-axis bend | 18.7px | 0.12px | 18.7px | 0.4px |
| Extreme 70° bend, close camera | 54.0px | 1.29px | 53.0px | 1.5px |
Read the table honestly: boundary error is fixed in every case (54px → 1.3px worst), interior error is largely fixed for fronto-parallel bends and barely improved under strong yaw. The yaw rows are the model's stated limit, not a bug to be found later.
Why not real markerless tracking
Unchanged from Phase 5's decision record, and for the same reasons: OpenCV via OpenCvSharp has no musl/Alpine binaries the WorkflowEngine image could consume without building it from source, which means hundreds of MB of image growth and a large native attack surface parsing untrusted media, and a from-scratch pyramidal Lucas-Kanade + RANSAC feature tracker could not honestly be claimed robust here. Nothing about extending the tracker to bent surfaces changes that calculus — if anything it strengthens it, since the work below shows how far a silhouette alone can be pushed. If true 3D pose (surface normal, depth, real foreshortening) is ever required, that does need a real CV dependency, and this design does not pretend otherwise. What it buys instead is the capability that a bent plate composites cleanly, with the constraint stated: the footage must contain a uniform-colour plate, and the interior mapping is an approximation.
Detection: ChromaPlateEdgeProfiler
Per frame, per plate, against the mask the tracker already built. For each of the four edges it
walks the edge in the quad's own projective parameter t — the same parameterization
InsertSurfaceMesh replays, so measurement and reconstruction cannot disagree — and at each
sample marches from a seed known to be inside the plate outward until the mask stops being set.
That gives a signed deviation profile in units of the edge's own "across the plate" vector:
dimensionless, so a number measured on the 320x180 tracking grid replays at 1080p with no
conversion.
The measurement is two passes, and real footage is why. The first version pinned the deviation profile to zero at both ends, deliberately refusing to move the tracker's corners so that curvature measurement stayed separable from corner-fitting accuracy. Run against the project's real clips that was simply wrong. Every real screen has ROUNDED corners, and the extreme-point corner fit lands on the rounding's 45° tangent point, several pixels inside where the straight edges meet; measured against those inset corners, a perfectly flat plate looks bowed on all four edges. The real phone clip read 0.027 of curvature against a 0.01 threshold, and a barrel-distorted version of the same clip read 0.025 — flat and curved indistinguishable, with the flat case over the line.
So: pass one fits a FREE quadratic to each edge over its middle band only (15% dropped at each end, exactly where rounding lives) and extrapolates back to the endpoints; each corner moves by the sum of the two contributions from the two edges meeting there, which are roughly orthogonal and so compose into a proper 2D correction. Pass two re-probes against the refined quad with the profile pinned at the ends again, and what survives is the genuine bend. A rounded-corner rectangle comes out "corners nudged outward, edges straight"; a curved plate comes out "corners roughly unchanged, edges bowed". The real clips now read 0.007 (flat) against 0.056 (barrel-curved).
The reading is grid-resolution dependent, so cross-check a near-threshold one
Measured while choosing footage for the alpha launch film, on two real unmodified clips (a phone
held in frame, and a desktop monitor showing a full green screen with hands typing in front of it),
by running the same analyze step twice and changing only InsertGridWidth/InsertGridHeight:
| Clip | Auto grid (0/0, ~320px long axis) | Forced 480x270 |
|---|---|---|
| desktop monitor | 0.0065 → Flat | 0.0179 → Curved |
| phone | 0.0056 → Flat | 0.0133 → Curved |
Both plates are physically flat screens, and both cross the 0.01 threshold purely by being probed
on a denser grid. The mechanism is the obvious one — a finer mask resolves more of the anti-aliased
edge, and the residual bow the refined-corner pass cannot absorb grows with it — but the practical
consequence is worth stating: a curvature reading just above MinInsertCurvature is not evidence
of physical bend. Re-probe it at a second grid resolution; a genuinely curved plate reads curved
at both, while a flat one flips. Auto is also simply the right default for a mixed
portrait/landscape source set, since it preserves each source's own aspect per
the tracking grid section.
The bow model is two parameters per edge — dev(t) = 4t(1-t)(bow + skew(2t-1)), a symmetric bulge
plus the leaning one a yawed bend produces. Deliberately not a free-form boundary: a free-form fit
would follow a phone's notch, a thumb crossing the bezel, or a compression artifact, and the Coons
blend would ripple that dent through the whole interior. The two-parameter family is structurally
the wrong shape to express a notch, so notches fall out as fit outliers instead of becoming
geometry — and it is four numbers per frame, which keeps the track temporally stable and the
artifact small.
Two validity gates, both forced by real footage
The project's laptop clip defeats the extreme-point corner fit outright: the plate is a strong trapezoid, "top-left" (min of x+y) lands partway down the left edge, and the fitted quad cuts the entire upper-left triangle off the plate — at confidence 1.000 and fill ratio 1.000, a bad-but-confident detection. The silhouette then genuinely does sit far outside that quad's top edge, and the profiler dutifully measured a 0.118 bow. Meshing on that reading made the composite visibly worse than the flat pin: it bent the top edge into a parabola that matches a triangle nowhere and crushed the interior. Two independent gates now catch it:
| Gate | Rejects when | Flat phone | Barrel-curved | Laptop (broken fit) |
|---|---|---|---|---|
| Corner validity | pass one wants to move a corner more than 0.10 of its across-vector — further than rounding could ever explain | 0.025 | 0.055 | 0.141 |
| Fit residual | the pinned profile's RMS residual exceeds 0.012, i.e. the model does not describe this boundary | 0.005 | 0.004 | 0.036 |
The residual tolerance is deliberately ABSOLUTE and not scaled by the bow's size: the residual floor is grid quantization, which does not grow with curvature — the barrel clip carries four times the bow at a lower residual than the flat one. Either gate firing discards the plate's curvature and keeps the tracker's own corners, so the insert falls back to the corner pin, which is no worse than before. The laptop clip now reads 0.000 and composites exactly as it did.
(The underlying corner-fit failure is a Phase 5 tracker bug, not one this addition introduces — noted here because it is what the gates exist to survive.)
Geometry: InsertSurfaceMesh
The interior is the projective map H of the four (refined) corners — identical to what the
single-quad perspective path already applies — plus a transfinite (Coons) blend of the four
measured boundary deviations:
P(u,v) = H(u,v)
+ (1-v)·devTop(u) · (H(u,1) - H(u,0))
+ v ·devBottom(u) · (H(u,0) - H(u,1))
+ (1-u)·devLeft(v) · (H(1,v) - H(0,v))
+ u ·devRight(v) · (H(0,v) - H(1,v))
Two properties make this the right base. All-flat bows reproduce the corner homography
exactly, at any mesh density: a projective map is uniquely determined by four point
correspondences, so a cell whose corners are sampled from H refits to H itself. A mesh over a
flat plate is therefore the same warp as one big perspective, which is what lets the executor
route flat plates to the original builder rather than approximating it. And every deviation
vanishes at the corners, so the Coons construction's four corner terms are identically zero and
the blend needs no corner-correction term.
Density is derived, not configured: a parabolic bow of peak s split into n cells leaves a
per-cell chord error of s/n², so n = ceil(sqrt(s / InsertMeshTolerancePx)) is the exact
subdivision that brings the straight cell edges within tolerance of the measured curve. Top/bottom
bows drive columns, left/right bows drive rows, and MaxInsertMeshCells caps the product. The
density is chosen once per insert from the MEDIAN bow across its keyframes, so one badly-read
frame cannot inflate the graph and each cell's canvas can stay fixed for the insert's duration.
Rendering: ScreenInsertMeshFilterBuilder
A single projective transform cannot represent a non-planar surface at all, so the surface is
tiled into Rows x Cols cells, each its own perspective warp.
remap was the alternative and was rejected on measured grounds. ffmpeg's remap takes an
arbitrary per-pixel coordinate map and would need no tiling whatsoever — but it samples
NEAREST-NEIGHBOUR from integer gray16 maps, which shimmers badly on exactly the content this
feature exists to composite (a UI with text, scaled down into a phone-sized region), and it would
require generating and decoding gigabytes of raw per-frame map streams: a new I/O mechanism, where
the tiled form stays what every other filter builder in this feature already is, a pure string.
The tiling costs encode time instead, and that was measured as cheap — a 3x6 mesh at 1080p ran at
94 fps against 253 fps for the single-quad path.
Per cell, on its own small canvas sized to that cell's destination bounding box over the whole insert (not the full frame — that is what keeps a dense mesh affordable and keeps each source cell resampled at roughly its destination scale):
[asset] fps, format, setpts(+outputStart), split into Rows*Cols
[cell] crop to this cell's source sub-rectangle (floor-tiled, so cells abut exactly),
scale to (cw-2)x(ch-2), pad to cw x ch with a 1px black border,
perspective sense=destination eval=frame -> warped cell
color=white (cw-2)x(ch-2), pad 1px black, format=gray, SAME perspective -> warped cell mask
alphamerge ; overlay at the cell canvas origin, enable='between(t,start,end)'
Seams. Two mechanisms, and the second was found by rendering real footage rather than reasoned
out. First, perspective is handed not the cell's mesh corners but the exact projective
EXTRAPOLATION of them out to the padded canvas corners — the content sits 1px inside the canvas
because of the border the warp needs, so asking it to send the CANVAS corners to the mesh points
would land the content a border-pixel short on every side and open a gap at every interior
boundary. Second, real renders still showed a green hairline along every interior boundary:
perspective interpolates bilinearly, so both the warped content and the warped white mask fade
out over roughly a pixel at the cell edge, and two abutting cells each contributing ~50% alpha
composite to ~75% coverage with the plate showing through the rest. Each cell is now grown 2px
onto every side that HAS a neighbour, so one cell's fully opaque interior covers its neighbour's
ramp. The mesh's outer boundary is never grown — it must stay exactly on the measured silhouette
(InsertOverscan remains the separate, deliberate knob for spilling past the plate edge).
Mesh and flat inserts interleave freely inside one chain: ScreenInsertFilterBuilder dispatches
per insert and the mesh builder ends at the same label the flat branch would have, so the compile
step's input bookkeeping, stage ordering and label plumbing are untouched. A mesh insert is still
ONE extra ffmpeg input however many cells it has.
Multi-plate ("multipoint") tracking
The interpretation pursued here is several independent chroma plates visible at the same time
— two devices in one shot, or a device and a separate monitor — each becoming its own r{n}
region with its own rendered asset. A single physical object with two rigidly-hinged faces (an
open laptop's screen and keyboard deck) falls out of the same mechanism as two plates; what is
NOT modelled is any constraint linking them, so their tracks are independent and could in
principle drift relative to each other.
ChromaQuadTracker now keeps the MaxPlatesPerFrame largest components clearing the area/fill
gates, in descending area order with ties broken by raster position, and associates them across
frames greedily by centroid distance — all candidate (track, detection) pairs in ascending
distance, ties broken by index, so the assembly is bit-for-bit reproducible. Tracks unmatched for
longer than MaxGapFrames close; unmatched detections open new ones. At the default cap of 1
there is at most one detection per frame and this reduces to the original single-run walk, with
one deliberate improvement: a track stays open across the gap window, so a plate that jumps away
and returns within it is re-associated instead of split.
The compile side needed no changes at all — ResolveInsertsAsync already looped over
plan.Inserts and chained one composite per region, so simultaneous regions "just work" once the
analyze step produces them.
Agent surface
Unchanged in shape. view.insertRegions gains one descriptive word, surface: "Flat" | "Curved",
so the planner can render content suited to a bent screen. The bow numbers, the refined corners
and the per-frame tracking data stay compile-time-only in the artifact — the same discipline as
every other id family. ScreenInsert is untouched: still exactly three strings, still pinned by
MotionGraphicsPlanOutputInvariantTests. A curved surface is a property of the SHOT, measured
by first-party C#; it is not something an agent may assert, request or influence.
Config reference
VideoAnalyzeStepConfig | Default | Meaning |
|---|---|---|
MaxInsertPlatesPerFrame | 1 | Chroma plates tracked simultaneously per frame. 1 is the original behaviour exactly. Clamped 1..16; MaxInsertRegions still caps the artifact total |
MeasureInsertCurvature | true | Measure refined corners + edge curvature. false reproduces the pre-existing composite exactly |
VideoCompileStepConfig | Default | Meaning |
|---|---|---|
InsertSurface | Auto | Auto meshes a plate whose measured curvature clears the threshold; Planar never meshes; Mesh always tries (debugging only — on a flat plate it is the same warp, more expensively) |
MinInsertCurvature | 0.01 | Curvature a plate must exceed under Auto to earn a mesh, as a fraction of its own size |
InsertMeshTolerancePx | 0.6 | Target worst-case chord error per cell; this is what sets the cell count |
MaxInsertMeshCells | 24 | Ceiling on rows * cols for one insert |
Real-footage results
Validated against the project's own green-screen clips (project
2f9216f9-dc7d-4118-9e60-e2855cb38218), driving the shipped ChromaQuadTracker →
VideoCompileStepExecutor → ScreenInsertFilterBuilder path, composited with a solid-colour asset
so "is this pixel content?" is exact. Registration is scored TWO-SIDED against the source frame's
own chroma mask, because green-leak alone rewards simply over-covering — which is precisely how a
flat pin can hide curvature:
- leak — plate pixel NOT covered by content: green showing through, the visible failure.
- spill — content pixel NOT on the plate: the insert painted onto the bezel or the hand.
Absolute spill is large for every variant and is not by itself error: InsertOverscan deliberately
expands the quad by 2%, and the source mask is conservative at the plate's anti-aliased edge. Only
the DIFFERENCES between variants mean anything.
| Clip | Variant | leak | spill |
|---|---|---|---|
greenscreen-phone.mp4 — flat plate, rounded corners + notch (reads 0.006, stays planar) | fitted corners | 0.55% | 39.7% |
| refined corners | 0.22% | 43.8% | |
| the same clip through a known barrel distortion (reads 0.033 → mesh 5x4) | fitted corners, flat pin | 0.24% | 41.1% |
| refined corners, flat pin | 0.16% | 46.5% | |
| refined corners, mesh warp | 0.18% | 41.3% | |
greenscreen-laptop.mp4 — flat plate, strong trapezoid (reads 0.015) | either path | ~0.00% | — |
| both real clips composited into one frame | two concurrent regions | both plates composited independently, each tracking its own motion |
Read plainly:
- Corner refinement is the clear win, and it is a trade. It more than halves leak on a real flat plate (0.55% → 0.22%) by pushing the quad out to where the straight edges actually meet, and pays for it in a little more content on the bezel. That is the right side of the trade — green showing through a screen is glaring, a few pixels over a dark bezel is not.
- On a curved plate the mesh buys back exactly that cost. At the same leak it drops spill from 46.5% to 41.3%, i.e. it CONFORMS to the plate where the flat pin over-covers it. Refined corners plus mesh is strictly better than the original on both axes (leak 0.24% → 0.18%, spill unchanged).
- The mesh does NOT reduce green on this footage, and it would be wrong to claim it does. With refined corners and 2% overscan the flat pin already covers a mildly curved plate; what the mesh improves is registration, not coverage. Its coverage advantage is demonstrated against analytically known geometry (the table further up: boundary error 54px → 1.3px), not against any clip in this project.
The barrel-distortion clip is the honest part of this to be clear about: it is real footage put through a known geometric distortion, which is what a wide-angle lens does to a flat plate, not a camera-captured curved object. It exercises the capability on real pixels with real motion, compression and noise, and a genuinely curved silhouette. Footage that does not exist in the project, and that would be needed to validate what is left: a physically curved screen or a plate on a cylindrical object — with enough bend that a flat pin cannot hide it under overscan — and a single shot containing two devices at once (the two-plate result above is a spatial composite of two real single-plate clips).
Sourcing a genuinely curved plate: searched, not found
Looked for while producing the alpha launch film, across Pexels, Pixabay, Mixkit, Videezy, Hollywood Camera Work's free VFX plates, and an ID sweep of one studio's fabric series, under every phrasing that seemed likely: green/blue/magenta cloth, fabric, flag, banner, curtain, tarpaulin, veil, drape, blanket, morphsuit, cylindrical product wraps, held cards and posters.
The finding is consistent enough to be worth writing down. In free stock libraries a chroma colour appears in exactly two forms: as a backdrop behind a subject, or as a full-frame fabric texture shot edge to edge. Both are useless here for the same reason — the tracker needs the plate's OUTLINE, and neither has one. The near misses illustrate it: studio fabric series shoot beautiful billowing cloth as an outlined foreground object against black, but in white, red, yellow and cream; the one magenta clip in that series has the magenta as the BACKDROP with a white cloth in front of it. Green morphsuit performers exist against black, but a head-plus-torso-plus-arms silhouette is not a quadrilateral and the validity gates reject it, correctly.
So a camera-captured curved chroma plate was NOT obtainable, and synthesising one — warping a flat clip to manufacture a passing curvature reading — would be fabricating the test data the gates exist to evaluate. If this capability is to be validated against real footage, the realistic route is to SHOOT it: a green cloth taped over a cylinder, or a sheet of green card bowed between two hands, is a few minutes of work with any phone and would settle the question properly.
Limits, stated
- Interior foreshortening is approximate under strong yaw or strong convexity — see the table above. The silhouette does not carry the information; only real 3D pose estimation would.
- A convex plate's corners are biased. When a plate bulges outward, the extreme-point corner fit lands on the bulge rather than at the true corner, and the bow is measured relative to that. The corner refinement absorbs part of it; the validity gate rejects the rest rather than guessing.
- Corner-fit failures are survived, not repaired. The laptop case degrades to the Phase 5 behaviour. Fixing the extreme-point fit itself is separate work.
- No inter-plate constraint. Two plates on one rigid object track independently.
- Everything Phase 5 deferred that is not listed above stays deferred: occlusion recovery beyond gap bridging, sub-pixel corner refinement at native resolution, inserts on multi-source or crossfade compiles, lighting/colour match of the insert to the scene, motion blur.
Silhouette matting
A tracked insert has to be cut to the shape of the plate it goes into, not to the quadrilateral that approximates it. Phase 5 and the non-planar work above both answer where the content goes; neither answers which pixels it covers, and until this section's work every composite was a filled quad. On the project's own footage that meant the phone's rounded bezel corners and its camera notch were painted over, and a hand crossing the laptop's screen was painted over too.
A real plate's silhouette is not a quadrilateral and cannot be made into one by bending its edges. It has to be cut out, per frame.
warp decides WHERE the content lands (quad / mesh, from tracked corners)
matte decides WHICH PIXELS it covers (this section)
spill suppression decides WHAT COLOUR survives (this section)
The cutout already exists in the source
The plate is a known colour, so keying the base picture gives the matte directly: pixel-exact, at full encode resolution rather than the coarse tracking grid, anti-aliased by the source's own sampling, and free of any extra decoding. Rounded corners, notches, an occluding hand and the motion blur on that hand all come along with it, because they are all simply "not the plate colour".
The insert's alpha therefore becomes a conjunction:
alpha = warped quad/mesh mask AND plate matte keyed from the base
Both halves are load-bearing. The quad mask keeps green scenery elsewhere in frame out; the plate matte keeps the bezel, the notch and anything in front of the plate in.
Why the key is searched rather than configured
A key needs a colour and a tolerance, and no fixed pair works. The project's two clips measure the
plate at RGB (0, 232, 0) — bright and saturated — and (31, 63, 10) — dark and nearly
neutral. A tolerance that keys the second keys half a room in the first. ffmpeg's chromakey,
which ignores luma and compares chroma only, fails outright on the dark plate for exactly that
reason: its chroma sits close to the rest of a dim room.
Rather than expose a knob nobody can set without looking at frames, PlateChromaKeySolver
grid-searches colour × similarity × blend per plate and scores each candidate against ground
truth it already has: the plate mask ChromaQuadTracker segmented for that very frame.
| Element | How it is chosen |
|---|---|
| Candidate colours | Percentiles of the plate's own luma distribution (25/40/50/60/75) — percentiles, not a mean, because a plate routinely carries a specular band and a dark fringe whose mean is a colour appearing nowhere on it |
| Similarity | Swept across the full usable range, finely enough that the winner is not a coarse compromise |
| Blend | Swept as fractions of the chosen similarity, so the edge ramp scales with the key's own tolerance |
| Score | Soft IoU (Σmin/Σmax) against the tracker's segmentation, plus a spill term (below) |
The distance metric replicates ffmpeg's vf_colorkey exactly — diff = sqrt((dr²+dg²+db²)/3)/255,
alpha 0 at diff ≤ similarity ramping to 1 by similarity + blend — verified against a real
ffmpeg run over a synthesized distance ramp, because a score computed against a different curve
than the one that will run would be worthless.
Candidates are judged only inside the plate's dilated bounding box. That is the only place the matte decides anything: the warped quad mask already excludes everything else, so a candidate that keys foliage across the room is not penalised for it, while one that eats into the bezel — the failure that matters — is.
Measured on the real clips, with no clip-specific constant anywhere:
| Clip | Solved key | Similarity | Agreement |
|---|---|---|---|
greenscreen-phone.mp4 | #00EC00 | 0.368 | 0.77 |
greenscreen-laptop.mp4 | #193E08 | 0.056 | 0.98 |
The spill term
Where something crosses in front of a plate, the pixels along its edge are a physical blend of that object and the plate behind it. They carry real green, more of it the more motion blur widens the edge. A key tight enough to match the tracker's segmentation leaves them in the base, and the composite then shows a coloured rim tracing the occluder.
Plain agreement cannot fix this: that rim is a handful of pixels against a whole plate and rounds away in the IoU. The objective therefore carries a second term that measures the miss normalised over the blend band itself, giving those pixels weight proportional to the problem they cause. The band is the difference between the tracker's own two detection thresholds — green enough for the permissive grow test, not for the strict core test — restricted to inside the tracked quad so that green scenery behind the plate can never drive the key wider.
Two deliberate details:
- Selection and reporting are different numbers. The spill term steers the choice; the stored
Scorestays plain agreement. Folded into one number, a plate with an inherently wide fringe would have its matte switched off for having a hard problem rather than for being badly solved. - The floor is lexicographic. Any candidate clearing
MinUsableScoreoutranks any candidate that does not, whatever their spill — otherwise a heavy spill penalty could hand the win to a key whose agreement had fallen under the floor, which reads downstream as "no matte at all".
A plate whose best candidate still disagrees with the tracker stores no key, and the compile step then composites with the plain quad mask exactly as it did before matting existed.
Spill suppression
Matting cannot fix a pixel it is correctly keeping. Once the key is as wide as it can go without swallowing the occluder, what remains has to be corrected photometrically:
G' = min(G, max(R, B)) the classic despill
= max(min(G,R), min(G,B)) the identity actually built, from two channel-mixed copies
and a darken/lighten pair — no plane surgery, R and B untouched
It is, by construction, a no-op on any pixel where green is not dominant, so it cannot
discolour skin, a bezel or a wall. Verified on real pixels: plate (28,66,10) → (28,28,10), skin
(145,103,69) → unchanged.
It must run in planar RGB, and for a long time it did not. blend has no packed-RGB support,
so ffmpeg converts its inputs to something it does support, and inside a full compile graph that
negotiation picked a YUV format, where darken and lighten are no longer per-RGB-channel min and
max. Measured on the phone clip: (33,73,32) came out as (6,50,5) rather than (33,33,32), so
suppression barely suppressed, and the rim it exists to clean stayed green. Run on its own, the
same chain happened to negotiate gbrp, which is why it looked right in isolation. Every branch
of it is now pinned with format=gbrp. Residual green round the phone insert fell from 61-113 px
a frame to zero on the colorkey matte path.
The corrected picture is laid over the base with the band as its alpha, not maskedmerged. The
band is synthesized a margin past the insert's window, and maskedmerge repeats every input's last
frame until the longest one ends (this ffmpeg has no shortest option for it). The finished film
therefore ran about a second past its source on its held final frame, after the insert's enable
window had closed, and any video whose insert reached its last frame ended on a raw green screen.
An overlay always ends with its main input, the base picture.
It is still bounded, because the one thing it would wrongly touch is genuinely green scenery. The bound is the insert's own warped footprint minus the matte — precisely the occluder where it crosses the plate, at whatever width its motion blur happens to be. Two other bounds were tried and are worth recording as dead ends: a fixed dilation radius is too narrow for a fast-moving edge and too wide for a locked-off one, and deriving that radius from the plate's blend statistics fails because the region those statistics describe (grow-minus-core) is the plate's dim interior, not its edge — it measured 145 pixels on a clip whose edge is one or two.
Overscan is inverted by matting
Without a matte, expanding the quad past the plate paints content onto the bezel, so
InsertOverscan defaults to a cautious 2%. That leaves the content stopping a hair short of the
plate wherever the coarse tracking grid rounded a corner inward — visible on real footage as a
green seam down the plate's left and bottom edges.
With a matte the content is clipped to the plate's silhouette however far it is expanded, so overscan becomes free and generous is strictly better: it guarantees the content reaches every pixel the matte will admit. A matted insert therefore uses at least 10%, never reducing a larger configured value. Measured: residual plate visible inside the composite fell from 2066 px to 545 px on the laptop occlusion frame.
How the content reaches that band changed later. It used to be padded by the same factor (black) and warped onto the expanded quad, which lands the content's own frame back on the plate only for an unforeshortened plate; a flat matted insert now pins the content's frame to the tracked plate itself and lets the warp smear its edge pixels across the band — see "Smeared overscan" in Tracking quality on the customer clips.
Sub-pixel corners at the source's frame rate
A tracked insert wobbled, visibly, even when its outline followed the plate well. On the project's generated phone clip (1344x768, 24 fps), measured against a full-resolution reference track built from edge-line fits, the corners the compile step consumed sat 3.0 px off the true corners on average and moved 3.7 px RMS (6.8 px at the 95th percentile) around that offset, every corner independently, which is what makes inserted content shear. Three separate causes, each fixed separately:
| Cause | Fix | Wobble after |
|---|---|---|
Each corner was ONE grid pixel (the extreme of x±y), quantized to ~4 source px and landing on the corner rounding, ~5 px inside the true corner | ChromaEdgeLineFitter: each edge located as a line through 40 sub-pixel crossings, corners where the lines meet | 1.9 px |
| The grid was sampled at 10 fps from a 24 fps source, so every sample was the NEAREST source frame, up to 21 ms off its stamped time | Sample at the source's own rate, or an exact divisor of it (InsertSampleFps = 0, auto) | 1.6 px |
| The default grid was 320 px wide | 480 px (ChromaQuadTracker.AutoGridBaseWidth) | 1.4 px |
The reference itself is noisy at about 1 px, so the result is close to what can be measured. The offset from the true corners fell from 3.0 to 1.0 px.
The edge fit. Along the middle 80% of each edge (corner rounding lives at the ends), a probe marches across the edge through the bilinearly interpolated dominance signal, target channel minus the strongest rival in absolute 0..255 units, and records where it falls through half the plate's own level. That 50% point of the soft plate/bezel transition is located from the intensity ramp to a fraction of a grid pixel. Absolute dominance rather than the brightness-normalised purity the growth bar uses, because the plate is usually framed by a near-black bezel, where purity is a ratio of small numbers and is noise: fitting to purity put every edge 3-5 px out on the bezel. A Tukey-weighted total-least-squares line through each edge's crossings ignores the stretch a reflection, a fingertip or a notch displaced, and a probe that does not start inside the plate is skipped rather than guessed. The fit returns nothing, and the extreme-point corners are kept, when an edge has too few crossings, a corner would move further than 12% of the diagonal, or the result is not convex.
Flat plates skip the profiler. The same fit measures how straight each edge is (the bulge of a
quadratic through its residuals, as a fraction of the plate's extent across it). A plate whose
edges are all straight (FlatEdgeBowFraction, 0.025) is fully described by its four line-fitted
corners, and those become its surface corners too. Previously ChromaPlateEdgeProfiler re-refined
every plate from the binary component mask, read grid ripple on the flat phone as a bow on most
frames, and moved the corners the compile step uses, adding its own ~1.5 px of wobble. Measured
on the phone: median 0.007, 90th percentile 0.011. A genuinely curved plate (the synthetic barrel
test reads 0.14) still goes through the profiler and keeps its mesh warp.
Frame-count knobs keep their meaning in seconds. MaxGapFrames, MaxOccludedFrames and the
occlusion and consistency windows were tuned at 10 fps; Track rescales them by
fps / ReferenceSampleFps, never below 1x. The 3-frame smoothing window is deliberately left in
frames: it averages per-frame measurement noise, and at a higher rate it lags real motion less.
Memory. The insert grid is read whole (the tuner and assembly both walk it), so the automatic
rate is chosen inside a fixed 640 MB budget (InsertGridByteBudget), about what the old
3000-frame cap cost at 320x180. A 6.6 s clip samples every frame at 480x274. A long clip drops to
the largest exact divisor that fits (a two-minute 24 fps clip samples at 12), and only when even
2 fps would not fit does it fall back to an arbitrary rate.
Dense keyframes reach the encode. MaxInsertExprKeyframes used to thin every track to 96
keyframes, roughly 4 fps on a 24 s insert, reintroducing the interpolation error dense sampling
removes. Past 96 keyframes the corner series is now written as a balanced tree of if(lt(in,…))
splits (log-depth nesting, which ffmpeg's expression parser accepts at any length), and the
default cap is 1500. Up to 96 keyframes the filtergraph is byte-identical to before.
A weak fragment at a track's end is dropped (ChromaQuadTracker.TrimWeakEnds). On the phone
clip the phone is turned toward the camera through a flash of glare: two frames see a thin sliver
of the screen edge-on, the next five show a pale reflection with no green at all, then the plate
is clean. Gap bridging joined the sliver to the body, so the insert opened on the sliver and
interpolated across the glare while the phone rotated: a quad sliding off the phone for a quarter
of a second. A markerless feature tracker (the Tracking/ LK + RANSAC pipeline) was tried for
those frames and loses the phone too, to fast rotation, motion blur and a screen whose appearance
flips from green to white. With no trustworthy geometry, the right composite is none. A leading or
trailing run of detections that is separated from the track's body by a bridged gap, and whose
median core area is under half the body's, is removed together with the bridge. A plate that
starts small and grows without a gap keeps every frame.
Tracking through a turn
After the sub-pixel work above the insert sat still on a still phone, and still wobbled while the phone was being TURNED toward the camera — the 20 frames of the project's phone clip where the plate is foreshortened, motion-blurred and crossed by reflection streaks. Three things were wrong, none of them "noise", and the first job was a measurement that could see them.
A reference for the turn. The earlier reference (edge-line fits to the green at full
resolution) broke on exactly the frames that mattered, because the green does. The reference used
here is bounded by the phone's dark BEZEL instead: from a temporal prior, 60 probes per edge march
outward through the full-resolution frame and take the dominance crossing where the probe starts on
plate colour, or the luminance drop into the bezel where it starts on a white reflection — the two
signals fail on different pixels, so together they cover the glare frames the green-only reference
could not ($S/gt_gen.py, propagated from the frame whose detection best agrees with its
neighbours). Its per-edge residual is 0.3 px and its second differences on the steady part imply
~0.4 px of noise. The metric reports, per corner, the error against it and the slip — the
frame-to-frame change of that error, which is content sliding on the glass and is what the eye
reads as wobble. Both were checked against overlays of the quads on the frames before anything was
optimised against them.
Diagnosis. Against that reference the per-frame detections were within ~1 px on EVERY frame of the turn except three: two where a reflection streak crossed the top and left edges and the edge-line fit followed the streak (63 and 41 px off), and the first frame of the track, where half the screen was glare and the fit bounded only the green half. The 3-frame average then spread each of those over its neighbours — a frame whose own detection was 0.9 px off came out 27 px off — which is why the insert looked wrong for six frames when three were measured wrong. Smoothing was the wrong tool; a corrupted measurement has to be recognised and replaced.
| Change | Why | Where |
|---|---|---|
| Corners in the plate's own frame | The extreme-point corner rule (TL = min x+y, …) assumes an upright plate; past ~30° two extremes converge on one physical corner and at 45° the "quad" is a triangle. Generated test shots turn a tablet from landscape to portrait and whip a phone up through 45°, and every frame of those turns fitted a kite. The orientation of the minimum-area rectangle around the component's convex hull is measured first and the rule applied to coordinates rotated by minus that angle; upright plates are unchanged | ChromaQuadOrientation |
| Labels follow the physical corners | That orientation is only defined modulo a quarter turn, so the fit renames the corners as a plate passes 45°, and content pinned to "top-left" would snap a quarter turn. Each matched detection is re-labelled by the cyclic shift closest to the track's previous sample (surface corners and edge bows shift with it). Continuity alone inherits whatever labelling the FIRST detection got, a coin toss when that frame sits near 45° — on the whip clip it put the content sideways on an upright phone for four seconds — so one further shift is applied to the whole track afterwards, the one that makes the "top" edge the most nearly horizontal over the track's median frame: the pose the plate holds longest is the one that reads upright, and a plate turning landscape → portrait gets whichever it held longer, continuously through the turn | ChromaQuadTracker.AlignCornerLabels, UprightTrackLabels |
| Implausible detections are replaced | Every detected sample is predicted by a straight line through its detected neighbours within ±125 ms; the furthest any corner sits from that prediction is compared with the larger of 2% of the plate's own diagonal and 1.5× the plate's own local per-frame travel. Both arms are relative to the footage — the floor scales with the plate, the speed arm with its motion — so a tablet filling the frame and a phone in the distance, a whip and a locked-off shot, are each judged against themselves, and a generated clip's frame-hold-then-double-step (one frame's travel off a straight line) passes. Rejection is greedy, worst first, re-predicting after each removal: judging every frame against predictions that still contain the outlier flagged all of the turn and then none of it. Rejected interior frames are interpolated from their clean neighbours; a rejected END frame is dropped. Capped at 30% of a track's detections | ChromaQuadTracker.RejectImplausibleSamples, Options.OutlierFloorFraction / OutlierSpeedFactor / MaxOutlierFraction |
| Smoothing is weighted by motion | The 3-frame average removes noise of the order of the measurement noise and nothing else; on a plate moving several pixels a frame it lags the motion by a fraction of a frame's travel and smears generated footage's frame holds. Each frame is blended between its own detection and the window's average with weight 1 / (1 + (motion / (2 × noise))²), motion being its neighbours' median per-frame travel and noise the track's own median residual from the rejection pass: a still plate is fully averaged, a plate moving at twice its noise half, a fast one passed through | ChromaQuadTracker.Smooth, Options.SmoothingMotionGate |
Measured (corner error / slip against the bezel reference, source pixels; $S/metric_gen.py).
Before is the tracker as shipped above; after is all four changes:
| Clip (generated with MiniMax, 24 fps) | Frames | Before: err mean / worst / slip RMS | After |
|---|---|---|---|
| Phone turned toward camera (1344x768), the turn, frames 26-45 | 20 | 4.46 / 35.2 / 7.01 | 1.48 / 9.5 / 2.72 |
| Same clip, steady part | 112 | 1.20 / 5.0 / 1.23 | 1.18 / 2.9 / 1.01 |
| Blue-screen phone tilted outdoors against foliage, handheld | 158 | 1.13 / 4.0 / 1.68 | 0.86 / 3.6 / 0.62 |
| Laptop on a desk, slow dolly, dim plate | 158 | 0.80 / 2.9 / 0.31 | 0.81 / 2.9 / 0.35 |
| Phone whipped up through 45° to fill the frame (768x1344) — frames where the plate is in picture | 98 | 1.09 / 12.1 / 0.76 | 0.99 / 3.7 / 1.04 |
| Tablet turned landscape → portrait by a window | 156 | 23.8 / 409 / 39.4 | 19.0 / 340 / 6.5 |
Variants tried on all six clips ($S/sweep_table.txt), each change isolated: the orientation fix
alone changes only the two turning clips (tablet slip 39 → 7); rejection alone takes the phone
turn from 4.46 to 2.46 px but leaves the average's lag (slip 4.6); the gate alone does nothing
measurable on the phone and is what fixes the blue clip; rejection without any smoothing is as
good on the turn (1.39) and worse everywhere still (phone steady 1.31 / 1.27, whip slip 1.74).
A hard on/off gate at 2× noise instead of the weight above was equal or worse on every clip (whip
slip 1.30 vs 1.04); 1× and 4× were each worse on at least two clips; a 5-frame window was worse
on the whip and the blue clip. Stricter rejection (1% / 1.0×) rejects frame holds and costs the
whip its worst frame back (11.9 px); laxer (4% / 2.5×) is identical on the phone and worse on
the tablet's slip. A finger-swipe clip whose plate runs off the top and bottom of the picture is
not in the table: the reference itself fails where the finger covers an edge, and that clip is
judged by eye (its right edge still follows the finger — occlusion of an EDGE, as opposed to the
area collapse "Surviving occlusion" detects, was an open problem here; it is fixed in
Tracking quality on the customer clips).
The two kite cases are where the orientation fix shows: the whip's worst frame goes from 12 to 4 px, and the tablet, which no longer fits a kite between 30° and 60°, drops its slip six-fold. The tablet's remaining error is a different problem, left open: the window reflection on its glass is green-TINTED, not white, so it sits below the purity bar the component grows through, the component excludes it, and the coarse right edge — and the line fit that starts from it — lands on the reflection's boundary inside the screen. The laptop is a plate that never moves, where the gate is never shut and the numbers are unchanged within the reference's own noise.
What did not help. Replacing the moving average with a local quadratic (Savitzky–Golay, 5 or 7 frames) was worse on the turn (2.2 px) than either the gated average or no smoothing at all, because the real motion is a staircase of held and doubled frames that no polynomial follows. Smoothing in homography space and a Kalman smoother were not tried once the diagnosis showed the error was three gross outliers and not noise: no filter of the measurements fixes a measurement that is of the wrong thing. Bridging a rejected frame with a quadratic through its ±3 clean neighbours instead of a line gained 0.08 px mean and 2 px on the worst frame of the turn, but a ±4-frame window and a least-squares line through the same neighbours were both WORSE than plain interpolation between the two nearest — too fragile to ship for the gain, so a rejected frame is bridged linearly like a detection dropout. The frame-hold cadence of generated footage is itself the floor on what interpolation can do: the plate stands still for a frame and then moves twice as far, and the 9.5 px worst frame after the change is a bridged frame on that cadence.
Reflections
A screen insert is only convincing if the glass still looks like glass. On the project's phone
clip, the phone catches a reflection as it turns. Before this, the composite showed two artefacts.
The reflection streak was left un-keyed, so the raw plate showed through the insert as pale green
blobs. And a potted plant behind the subject was keyed as plate, so content was painted onto the
plant wherever the insert's overscan reached it. Both came from the colorkey matte, and the fix
has two parts: a matte keyed on the plate's own chroma level, and carrying the reflection over.
The dominance matte
colorkey keys by RGB distance from one colour, and that radius is the wrong shape on both sides.
A reflection is the plate plus white light, (150,215,160) against a plate of (0,202,0), far
outside any radius that excludes the room. Dull scenery green, (70,93,26), is close in RGB
distance to a dark plate. The tracker now records each plate's own median dominance
(VideoInsertChromaKey.PlateDominance, 0..255), and when it is present the compile step keys on
that instead. The matte is the maximum of three terms, each a function only of the pixel's
dominance x and its target-channel level y, so the whole matte is ONE lut2 table lookup:
| Term | Ramps over | Catches |
|---|---|---|
| Plate | dominance 35%..65% of the plate's level | the plate; the ramp is centred on the same 50% point the edge fit puts the edge at |
| Dark plate | purity x/y 0.5..0.75, above an absolute floor of 10..24 | the shadowed plate and its 1-2 px anti-aliased edge into the bezel, (0,25,0): purity 1.0, a tenth of the plate's dominance. Without it that rim showed as a green outline |
| Reflection | dominance 8..24 AND target level 45%..65% of the plate's | glare on the plate: light added keeps the target channel near the plate's own (dimmest streak pixels ~121 against a plate of 202), dull scenery sits far below (the plant: 93) |
White walls and the pale blue shirt have no dominance at all and never key. The plant reads about
5% at the foot of the reflection ramp, the price of keeping the dimmest real streak pixels. An
artifact written before PlateDominance existed keeps its colorkey matte unchanged.
Carrying the reflection over
With InsertReflections on (the default; it needs a matte), the plate's reflection layer is
screen-blended over the composited insert, inside the matte and the insert's footprint only.
Glass shows what is behind it plus what it reflects, and screen is the clipped form of that
addition. The reflection layer is the ACHROMATIC part of the base picture, min(R,G,B) in every
channel less a noise floor of 8. A clean plate of any hue contributes black, a no-op under screen,
and white glare carries over intact. The plate's colour cannot simply be despilled away instead,
because real plates are not pure: the phone's reads (0,171,30) where the glass turns from the
light, despilling keeps that blue as "reflection", and over the insert's black content it tinted
the whole screen dark teal. The cost of taking only the neutral part is that a coloured
reflection, say a red shirt mirrored in the glass, carries over as grey glare.
Two ffmpeg details are load-bearing. The blend runs in planar RGB (gbrp): screen applied to YUV
planes blends the chroma planes too and shifts every colour of the insert (red read as pink). The
result is merged back with alphamerge + overlay, not maskedmerge: the screened copy differs
from the composite over the WHOLE frame, and maskedmerge, given it with this mask, returned the
screened copy everywhere, tinting the finished frame magenta edge to edge.
Surviving occlusion
Occlusion is missing data, not a smaller plate, and treating it as the latter is what made the tracker twitch.
When a hand crosses a screen the detector does not fail — it succeeds, and finds the remaining visible fragment. Fitting a quad to that fragment collapses the tracked geometry onto it, so the composited insert shrinks and snaps back as the hand passes. On the project's laptop clip that left a wedge of raw plate exposed for the whole pass, and it is not something the matte can repair: the matte can only clip content that the warp actually placed there.
A matched detection is therefore checked against the track's own recent behaviour, and one whose plate area has collapsed is discarded rather than fitted. The track holds its geometry, stays open, and the skipped frames are bridged by the same linear interpolation that already covered detection dropouts.
Two expiries, because two different things are happening
| Expiry | Governs | Default |
|---|---|---|
MaxGapFrames | A plate that has genuinely VANISHED — no detection at all | 3 |
MaxOccludedFrames | A plate still detected every frame, just partially | 45 |
The distinction is available for free: an occluded plate is still producing a detection. Past
MaxOccludedFrames the smaller reading is more likely to be what the plate now is than a very
long occlusion, so it is accepted and the baseline re-establishes — which avoids both holding a
stale geometry forever and splitting the track mid-shot.
Matching survives the hold
A track held through an occlusion goes stale as a match target. Matching therefore runs against a constant-velocity prediction of the track's centroid, with a jump allowance that grows with the number of frames held. Without this a plate occluded while moving falls outside the centroid gate and is adopted by a brand-new track — which is exactly how the first version of this fix cut a handheld phone's single track in two.
Three things that had to be measured, and were wrong at first
This is the part worth reading before changing any of it. Each of these was a plausible-looking choice that real footage falsified.
1. The test reads the STRICT CORE area, not the grown component. Hysteresis growth (see Tracked screen inserts (Phase 5)) intermittently bridges a plate into green scenery behind it. On the phone clip that makes the grown area flip between 0.19 and 0.35 frame to frame, while the core drifts smoothly by under 2%. Any threshold on the grown area is deaf or hysterical; the core is a steady signal.
2. The threshold is derived per track, not fixed. What counts as an abnormal dip is a property
of the footage. Against a rolling baseline, the unoccluded handheld phone never drops below 0.980,
while a hand crossing the laptop's screen takes it to 0.814 — but a locked-off shot of a bright
plate holds to a fraction of a percent, and a fixed ratio suiting one is wrong for the others.
Each track measures its own robust spread (median absolute relative step — median, so the very
dips being hunted cannot inflate the yardstick used to find them) and sets its threshold at
OcclusionSensitivity deviations below it, bounded by MinOcclusionDropMargin /
MaxOcclusionDropMargin.
3. The baseline is a THREE-sample median, not ten. A longer window lags a plate genuinely moving away from camera until the lag itself looks like a collapse — a steadily shrinking plate had its track cut short until this was narrowed. Three is enough to shrug off a single bad frame, which is all the baseline needs; the spread is still estimated over a longer window, where more samples genuinely help.
Asymmetry is the whole trick
Every part of this test is one-sided, and deliberately so: occlusion can only ever remove plate area. A plate moving toward the camera grows and is never suspected. One moving away shrinks gradually, which the short rolling baseline follows. Only an abrupt one-sided collapse looks like something passing in front.
Config reference
VideoAnalyzeStepConfig | Default | Meaning |
|---|---|---|
SolveInsertChromaKey | true | Grid-search each plate's colorkey parameters. Runs over already-sampled grid frames and is scored against the tracker's own segmentation, so it costs no extra decoding and needs no clip-specific configuration |
VideoCompileStepConfig | Default | Meaning |
|---|---|---|
EnableInsertMatte | true | Clip each insert to the plate's per-frame silhouette. A track with no solved key (or one that scored too low) composites with the plain quad mask regardless, so this can only improve a composite or leave it unchanged |
ChromaQuadTracker.Options | Default | Meaning |
|---|---|---|
OcclusionSensitivity | 6.0 | Robust deviations below a track's own behaviour before a core-area dip reads as occlusion |
MinOcclusionDropMargin | 0.03 | Smallest dip that may ever count as occlusion — below this is sensor and compression noise |
MaxOcclusionDropMargin | 0.30 | Largest dip a track may demand before believing an occlusion |
MaxOccludedFrames | 45 | How long geometry is held before the smaller reading is accepted as the truth |
Real-footage results
| Case | Before | After |
|---|---|---|
| Phone's rounded bezel corners and camera notch | painted over by a filled quad | followed exactly; notch cut out |
| Hand crossing the laptop screen | painted over; insert collapsed onto the visible fragment, exposing a wedge of plate | hand passes IN FRONT; geometry held through the pass |
| Residual plate inside the composite (laptop occlusion frame) | 2066 px | 545 px |
| Green rim tracing the hand | present | removed by the spill term plus despill; a trace remains at the blend limit |
Limits, stated
- The blend limit is physical. A pixel that is genuinely half plate and half occluder cannot be made entirely correct; it can be absorbed into the matte or despilled, and both are done, but a faint trace can remain on a fast-moving edge.
- The key is solved on ungraded pixels but applied to the base at the insert stage. A colour
grade applied earlier in the chain shifts the plate's colour out from under it.
EnableColorGradeis off by default, so this is latent rather than live; keying the original input instead is not possible because only the base is time-aligned with the output after the cut. - Hysteresis growth bridging a plate into green scenery was only worked around here at first; it is now fixed at the source by purity-gated, self-tuned growth — see Self-tuned growth and honest confidence.
- Occlusion recovery is a hold, not an estimate. Geometry is interpolated across the occluded frames from the good frames either side; a plate that moves non-linearly while fully occluded will be slightly wrong in the middle of the pass.
Self-tuned growth and honest confidence
What went wrong on the handheld phone clip
A tracking test on greenscreen-phone.mp4 — a phone held almost still in front of a lawn —
composited an insert that swung far outside the phone, reported Moving motion and 0.71
confidence, and painted over the phone's rounded corners. One cause produced all three:
- The permissive grow mask is a fixed "green dominates by ~10%" test. The lawn reads about
(70, 93, 26)and passes it. Wherever the coarse grid thinned the bezel to nothing, hysteresis growth flooded the plate ((0, 230, 0)) into the grass, and the extreme-point corners jumped onto the lawn on roughly half the frames — bottom-right corner at(0.68, 0.68)on one frame,(0.99, 0.97)on the next. - Every confidence term was computed one frame at a time, and each leaked frame was a tidy, well-filled quad, so nothing in the score could see the flicker.
- The chroma-key solver scores candidate keys against the tracker's own segmentation. Half of that segmentation was lawn, so no key agreed, none was stored, and the compile step fell back to the plain quad mask — no silhouette matte, rounded corners painted over.
- Separately, motion was the frame-to-frame centroid path length; one-grid-pixel corner jitter at
12 fps alone "travels" twice the
Movingbar.
Purity-gated growth
Growth now admits a pixel only if its chroma purity — how far the target channel rises above
the strongest other channel, relative to the target, (g − max(r, b)) / g for green — reaches a
fraction of the seed's own median purity. Purity is brightness-independent, so a dim region of a
plate scores like its bright regions while dull foliage scores far lower. Measured on the real
fixtures: the phone's lawn reads purity 0.1–0.3 against a plate of 1.0; the dim laptop plate's core
reads 0.63 and its genuinely dim regions 0.5–0.6.
Searching the bar per video (ChromaTrackerTuner)
What fraction of the seed's purity to require is a property of the footage, so it is not a
constant: before tracking, ChromaTrackerTuner tries 1.5 (effectively no growth), 0.8, 0.65,
0.5, 0.4, 0.3, 0.2 and 0 on up to four blocks of 18 consecutive frames spread across the
clip, and keeps the best by
score = detectionRate × edgeContrast × temporalConsistency
- edgeContrast — along each fitted edge (10%–90% of its length, so a rounded corner's arc does not count against it), the purest pixel within 4 grid px inside minus the least pure within 4 px outside, normalized by the purity at the quad's centre. Only a fit whose edges sit on the real plate boundary has contrast across them: a fit that leaked into scenery has scenery on both sides of its far edge, one that stopped short inside a dim plate has plate on both sides.
- temporalConsistency — the share of frames whose corners all sit within 4% of the quad's diagonal of their ±6-frame temporal median.
Ties go to the more conservative candidate. Measured on the five real fixtures:
| Clip | Old growth (bar 0) | Under-growth (bar 1.5) | Chosen bar | Chosen score | Track confidence |
|---|---|---|---|---|---|
greenscreen-phone.mp4 (lawn) | 0.42 (consistency 0.65) | 0.79 | 0.4 | 0.86 | 0.80 (was 0.71, wrong) |
greenscreen-laptop.mp4 (dim) | 0.89 | 0.47 | 0.4 | 0.89 | 0.92 (was 0.64) |
greenscreen-monitor-typing.mp4 | 1.00 | 1.00 | 1.5 | 1.00 | 0.97 |
c1-laptop-barrel30.mp4 | 0.85 | 0.44 | 0.5 | 0.87 | 0.91 |
c2-laptop-barrel48.mp4 | 0.77 | 0.47 | 0.5 | 0.85 | 0.87 |
The winning parameters and every candidate's score and measurements are stored on the track as
VideoInsertRegionTrack.Tuning, carried into the timeline manifest, and shown in the editor, so a
track can be explained from the artifact alone. ChromaQuadTracker.Options.GrowExcessFraction
pins the bar and skips the search.
On the phone clip the result is a plate that stays on the screen (corners move by about 1% of the
frame between frames instead of 30%), motion Slow, and a solved key (#00E200, 97% agreement),
so the insert is matted to the phone's rounded corners and notch again.
Honest confidence
Confidence is coverage × meanFillScore × meanEdgeContrast × consistency. The two new terms are
what can see a wrong fit: with the bar pinned to the old behaviour the phone track now reads below
the MinInsertConfidence floor of 0.5 and is refused rather than composited, while the correctly
fitted dim laptop plates are no longer docked for needing growth.
Motion is classified from the smoothed track over half-second strides, where back-and-forth jitter cancels and genuine travel accumulates.
Tracking quality on the customer clips
A customer's composited reels showed three things: the tracked screen jittered (worst on the finger-swipe phone clip and the turning tablet), raw green flashed through on some frames — a wedge down the phone while a finger crossed it at ~0:11, another down the tablet mid-turn at ~0:16, and the reel's last five seconds entirely green — and on a laptop/monitor reel the inserted UI was cropped at the plate's edges ("Dashboard" read "shboard"). Five separate causes, each measured and fixed separately; one of them was not a tracking problem at all.
The instrument
InsertTrackingQualityHarness (tests project) runs the analyze step's own tracking pass over a
directory of clips, composites a coordinate card into every track with the compile step's own
filter builder and real ffmpeg, and reports per clip. It is skipped unless
REELBOLT_INSERT_CLIPS names a directory; REELBOLT_INSERT_CLIP_LIST, REELBOLT_INSERT_LABEL,
REELBOLT_INSERT_BASELINE=1 (every change below switched off where it can be),
REELBOLT_INSERT_OPTS=Name=value,... (tracker option overrides, for sweeps),
REELBOLT_INSERT_SAVE_FRAME=clip:frame,..., REELBOLT_INSERT_CONTENT/REELBOLT_INSERT_FIT (real
content instead of the card) and REELBOLT_INSERT_DEBUG=1 (which detections the outlier pass
replaces) steer it.
| Metric | What it measures |
|---|---|
| Residual green | Per output frame, the strict-plate-green pixels left in the composite as a fraction of those in the source frame; frames over 2% are a visible flash, over 50% a raw plate. A "visible green" variant also counts the dim, reflection-covered plate the strict test misses |
| Reference error / slip | Every tracked frame's edges re-fitted on the SOURCE frame at native resolution, seeded with the track's own corners, on frames where all four reference edges sit on a clean plate/bezel boundary (edge contrast ≥ 0.6). Error is the corner distance to that reference; slip is its frame-to-frame change — content sliding on the glass, which is what the eye reads as jitter, and which also exposes smoothing lag on fast moves |
| Content extent | The card encodes its own x in red and y in blue, so the visible content's range says how much of its frame survives on the plate |
| Content-frame offset | Where the content's own corners land relative to the tracked plate under the old padded overscan (below) |
A quadratic-residual "jitter" was tried first and rejected as a metric: generated footage moves in a staircase of held and doubled frames, and the content SHOULD follow that staircase — the whip clip read 4.25 px of "jitter" against 1.05 px of slip.
Cause 1: every insert ran one frame ahead of its plate
perspective's in variable is 1-based — ffmpeg evaluates it as the link's frame count plus
one — and the corner expressions were written with keyframe 0 at in = 0. So every insert used the
NEXT frame's geometry. On a still plate that is invisible; on a moving one it is one frame's travel
of misregistration, a strip of raw plate along the trailing edge and content sliding against the
glass on every fast move. Verified directly (Perspective_in_variable_is_one_based...: a
keyframe-0 edge at x=100 rendered at x=150, half way to keyframe 1). Keyframes are now written at
in = k + 1 (ScreenInsertFilterBuilder.PerspectiveFirstFrame, both the flat and the mesh
builder). On the whip clip this alone took residual green from p95 3.76% to 0.07% and its
flashing frames from 10 to 5 (the rest are the whip's own first frames, see "What remains").
Cause 2: a finger over a corner, a reflection along an edge
The per-frame edge-line fit is seeded from the plate's extreme-point corners, and a COVERED corner is not where the plate's corner is. On the swipe clip a fingertip over the lower-right corner dragged that corner up to 40 px along the finger; the right edge was fitted between the true top-right corner and the fingertip, and a green wedge opened for two thirds of a second. The tablet did the same along the boundary of a green-tinted window reflection on its glass. The occlusion test in track assembly cannot see either: the core area falls only a few percent while a finger enters, and a reflection costs no area at all.
ChromaEdgeTemporalRefiner re-locates every frame's edges from its NEIGHBOUR's. A plate moves
little between frames, so the previous frame's edge is a far better seed than this frame's own
corners. Per edge, candidates are scored by ChromaEdgeLineFitter.EdgeContrastScore — dominance
just inside the line minus just outside it, over the plate's level, along its middle 80%: a real
screen edge has plate inside and a near-black bezel outside and scores near 1; a line on a finger
has plate on both sides along most of its length, and a line along the tablet's reflection has
dim plate (dominance ~25 against a level of ~140) outside it.
| Candidate | Why it is offered |
|---|---|
| The detected edge | Wins every tie (PreferenceMargin 0.05), so a clean frame is left exactly as it was |
| A re-fit seeded on the neighbour's edge | Probes that start on the finger or the reflection are skipped as holes; the robust fit ignores the minority that cross the wrong boundary |
| On a frame bridged across an occlusion: the detection the occlusion test set aside, and a re-fit seeded on it | Its uncovered edges are still the plate's real edges (FrameQuad.SetAside) |
Two rules keep it from undoing what the occlusion hold protects. A candidate may not move an edge
net inward by more than 5% of the plate's extent across it (InwardToleranceFraction):
something in front of a plate can only remove plate area, so a line further in is an occluder's
boundary — "net", because the corruption repaired here TILTS an edge, inward at one end and outward
at the other. And corners move at most 35% of the diagonal (the reflection needed 21%). The pass
runs forward then backward, so a track whose first frames are the corrupted ones is repaired from
the clean frames after them; it covers flat tracks only, and treats a frame of a track that reads
flat overall (median curvature under the compile step's 0.01) as flat even when its own profile
read a bow (motion blur).
Cause 3: plates with nothing composited into them
Three ways a plate ends up on screen with no insert, all of them seen on the customer's reels:
- The track was dropped. The reel's last five seconds were not an untracked clip: its fourth
phone's track existed (confidence 0.62) and was dropped
max_inserts_exceeded—MaxInsertsdefaults to 3 and the reel had four clips. - Frames the tracker deliberately leaves out — a phone turning through glare
(
TrimWeakEnds), the first frames of a whip into place, a plate seen too briefly to become a track. On the man-with-phone clip that is frames 19-21 and 25, at full raw green. - Output time outside the insert's window, e.g. a track cut into two kept portions.
The analyze step now records, per track, UncoveredPlateSpans: every run of frames where the
plate is visible (strict plate pixels covering a quarter of MinInsertRegionAreaRatio) but no
track covers it, attached to the nearest track in time, with the bounding box of the plate pixels
seen over the run. The compile step's InsertPlateGuardPlanner turns every such window — and
every window of a track that is not being filled — into a plate guard: an ordinary insert
whose content is a dark neutral screen (0x101010, an in-graph lavfi colour), matted with the
track's own solved key exactly as a real insert is, so it can only darken pixels that ARE the plate,
with the plate's reflections carried over so it reads as a screen that is off. Geometry is the
track's own quads inside its span, and the recorded box (grown 10% a side) over an uncovered span.
A guard never overlaps the insert it accompanies. Each window is reported in the EDL
(inserts.unfilledPlates: region, reason, output window, neutralized) and the compile adds one
customer-facing C1 warning, insert_untracked_plate_visible. A track with no solved key, or far
under the confidence floor (below half of MinInsertConfidence — it may not be a plate at all), is
reported but not darkened.
To fill the reel's fourth phone rather than darken it, raise MaxInserts to 4.
Cause 4: the cropped laptop title was the configured fit
The "Dashboard" → "shboard" reel was compiled with insertFit: "Cover", not Auto: Cover crops the
content to the plate's shape, and the laptop plate's aspect is 1.40 against the UI's 1.78, so
10.5% came off each side — the sidebar and the start of the title. Re-rendered with the same
tracks and content: Cover reproduces the crop exactly; Auto contains it (whole frame, dark bars top
and bottom); Stretch fills the plate with the content squeezed 21% horizontally. Evidence crops are
kept in docs/screenshots/tracking-2026-10/ (laptop-as-delivered-title.png,
laptop-new-cover.png, laptop-new-auto.png, laptop-new-stretch.png). Use Auto or Contain for
content that must be seen whole.
The overscan was a suspect too, and measured a real but smaller error. A matted insert's warp
target is the plate grown by 10% about its vertex centroid, and the content used to be padded by
the same factor (black) so its own frame landed back on the plate — exactly, only for a
parallelogram. On a foreshortened plate the content's corners landed off the plate's, by a mean
2.8 px (max 4.9) on the laptop, 2.4 px (max 7.3) on the tablet, and up to 86 px on the whip's
degenerate first quads — a strip of black pad inside the plate, or content cropped. Smeared
overscan (ResolvedScreenInsert.ContentKeyframes) replaces it for flat matted inserts: the
content's frame is warped onto the tracked plate itself, without the 1px border, and
perspective's edge clamping smears its outermost pixels across the band, where the matte keeps
them to the silhouette. The content's corners are on the plate's by construction, and the band
between the tracked edge and the true edge shows the content's own edge colour instead of black. A
mesh insert keeps the pad.
Cause 5: smoothing — measured, left alone
With the corrupted frames repaired, the remaining slip on the moving clips (0.5-1.05 px) is near the reference's own noise. A sweep of the smoothing found no setting better on every clip, so the defaults stand:
| Setting | swipe | tablet | whip | man-phone | phone |
|---|---|---|---|---|---|
| Shipped: window 3, motion gate 2× noise | 0.90 | 0.54 | 1.05 | 1.05 | 0.37 |
| Gate 1× | 1.18 | 0.58 | 1.42 | 1.10 | 0.32 |
| Gate 4× | 0.78 | 0.55 | 0.80 | 1.24 | 0.41 |
| Window 5 | 0.79 | 0.64 | 1.11 | 0.94 | 0.50 |
| No smoothing | 1.43 | 0.68 | 1.76 | 1.27 | 0.29 |
(slip RMS, source px.) Disabling the outlier rejection was also tried: it lowered slip on the swipe (0.85) and whip (0.91) clips — the rejection does smooth away some real frame-hold cadence — but let three swipe frames flash green (up to 6.7%), so it stays.
Before and after
Before is the shipped tracker and compile; after is all of the above. Residual green: mean / p95 / frames over 2% / frames over 50%. Reference: mean / p95 / max corner error, and slip RMS, in source pixels.
| Clip | Residual before | Residual after | Reference before | Reference after |
|---|---|---|---|---|
| Finger swipe (1344x768) | 0.61% / 5.31% / 16 / 0 | 0.02% / 0.07% / 0 / 0 | 2.28 / 9.35 / 48.0, slip 1.67 | 0.78 / 1.94 / 3.8, slip 0.90 |
| Tablet turning by a window | 0.39% / 2.81% / 12 / 0 | 0.02% / 0.18% / 0 / 0 | 0.77 / 1.81 / 2.6, slip 0.54 ¹ | 0.77 / 1.81 / 2.6, slip 0.54 ¹ |
| Phone whipped up (768x1344) | 2.10% / 3.76% / 10 / 1 | 1.82% / 0.01% / 5 / 1 ² | 0.78 / 1.89 / 4.3, slip 1.05 | 0.78 / 1.89 / 4.3, slip 1.05 |
| Man turning a phone toward camera | 2.95% / 0.07% / 4 / 4 | 0.44% / 0.00% / 1 / 1 ³ | 0.98 / 1.93 / 2.7, slip 1.05 | 0.98 / 1.93 / 2.7, slip 1.05 |
greenscreen-phone.mp4 (control) | 0.26% / 0.57% / 0 / 0 | 0.26% / 0.57% / 0 / 0 | 0.70 / 1.24 / 2.1, slip 0.37 | unchanged |
| Laptop on a desk (control) | 0.00% / 0.00% / 0 / 0 | 0.00% / 0.00% / 0 / 0 | 0.60 / 1.83 / 2.5, slip 0.38 | unchanged |
| Monitor (control) | 0.00% / 0.00% / 0 / 0 | 0.00% / 0.00% / 0 / 0 | 0.67 / 0.99 / 1.0, slip 0.01 | unchanged |
¹ The reference itself fails on the tablet's reflection frames, so it covers only the 50 clean
ones, where nothing changed; the residual column is where the reflection fix shows. ² The five
frames left are the whip's first, see below. ³ One frame, see below. Before/after composites of the
swipe (frame 116), the tablet (frame 101) and the man-with-phone turn (frame 20) are in
docs/screenshots/tracking-2026-10/.
What remains
- The first ~0.3 s of a whip into place. Frames 40-46 of the whip clip — the phone at 45°, heavily motion-blurred, entering the frame — still fit a wrong quad, and 23-44% of the plate shows on four of them. The refiner cannot help (its seed is the equally wrong neighbour), and the guard does not overlap a running insert.
- A plate washed out by glare keys only partially, so a guard darkens it only partially (man-with-phone frame 25, 59% left).
- A plate in a clip with no track at all (every appearance shorter than
MinInsertRegionSeconds) has no key and no geometry to guard with, and still shows green. - Cover crops by design.
Memory: every branch starts at zero
A four-clip landscape compile with tracked screen inserts was SIGKILLed (ffmpeg encode failed (exitCode=137)) on a host with 44 GB free and no container limit: ffmpeg itself grew to tens of
gigabytes. The cause was the shape of a matted insert's filtergraph, not the number of clips.
Why it buffered. A matted insert's alpha is keyed off the base picture (see
Silhouette matting), so the base is one split with a branch per matte,
despill and reflection layer. The insert's content and quad-mask branches used to be moved onto the
output clock with a timestamp shift, setpts=PTS-STARTPTS+start/TB, so their FIRST frame sat at
the insert's start. Every framesync filter (overlay, alphamerge, blend) must see a frame on
each input before it can emit anything, so the overlay could not pass output frame 0 until the
insert branch produced its first frame — and that frame needed the matte at start, which needed
the base decoded up to start. split hands every frame to every output whether or not it is
consumed, so the whole programme before the insert's start was queued, full resolution, in RGB,
once per split branch. Memory grew with how LATE an insert started, times the number of matted
branches; a four-clip composite's last insert starts 20 s in.
The fix (ScreenInsertFilterBuilder.DelayToOutputStart). A matted insert's branches are no
longer shifted; they are PADDED from t=0 with black frames (tpad=start=N:color=black, N = the
shift in whole frames), which tpad generates without pulling its input. Every branch now carries
a frame for every output frame and advances in lockstep with the base. The pad is black, which on
the mask is zero coverage, and the overlay's enable window gates it anyway; the pad follows
perspective, so its per-frame in still counts from the insert's first frame; the mesh path pads
each cell after its own warp. Unmatted inserts keep the shift — nothing they wait on comes from the
base. A split of one content file for several plates was deliberately NOT introduced: split
branches consumed at different times are exactly what buffers, and a separate -i per insert costs
only a decoder.
Compile (1920x1080, preset slow) | Peak RSS before | Peak RSS after |
|---|---|---|
| Four clips, 4 inserts + 3 plate guards | > 20 GB (killed by the measuring harness; SIGKILLed in production) | 7.6 GB |
| Two clips (man with phone, finger swipe), 2 inserts + 2 guards | 14.7 GB | 4.7 GB |
Same two-clip graph without the encoder (framemd5) | 13.6 GB | 4.1 GB |
The two-clip graph's decoded output is bit-identical before and after (all 733 video and audio
framemd5 lines). ScreenInsertMemoryShapeTests pins the shape (no matted branch may contain a
PTS-STARTPTS+ shift; every one is padded after its warp) and measures a small real-ffmpeg case:
a matted insert 20 s into a 640x360 programme peaks at ~190 MB padded, ~900 MB shifted.
What memory still scales with. Not with how late an insert starts any more, but still with the number of matted inserts and plate guards: each is a few dozen full-resolution filter links (matte, despill band, reflection layer, warp, composite), and ffmpeg keeps a small pool of frames per link — about 0.6-0.9 GB per matted insert at 1080p (measured: 0.4 GB for the bare four-clip concat, 3.9 GB with four matted chains, 6.75 GB with seven). Restricting each chain to its plate's bounding box would shrink that roughly by the box's share of the frame; it is not done.
Measuring a compile. InsertCompileMemoryHarness (tests project, skipped unless
REELBOLT_INSERT_MEMORY_DIR names a directory holding the analysis artifact as analysis.json,
each source as <projectFileId>.mp4 and the insert content as content.mp4) runs the real
VideoCompileStepExecutor against them and, instead of encoding, copies the scratch space and the
encode's argv into capture-<label>/ so the exact ffmpeg invocation can be run under
/usr/bin/time -v. REELBOLT_INSERT_MEMORY_SOURCES=0,1 compiles a subset of the clips.
A clip's last frame and the next clip's first
On a multi-clip compile with inserts, the FIRST frame of every clip after the first showed its raw green plate — neither the insert nor the dark plate guard covered it (16% and 21% of the frame on the customer's two-clip composites); single-clip compiles were clean.
Generated clips carry fewer video frames than their container says: 158 frames at 24 fps (6.583 s)
in a 6.592 s file, the audio being longer. A span kept to a clip's end is snapped to the
container's duration — 159 frames, 6.625 s — so the output timeline put the next clip at 6.625 s.
The segment delivered only 158 frames, concat started the next clip at the audio's 6.592 s, and
that first frame fell before every insert's and guard's enable='between(t,6.625,...)' window.
Worse, every later frame of that clip paired with the previous frame's geometry. The timing in the
filters was right; the picture disagreed with the timeline.
Each segment of the segmented (concat) encode is now held to EXACTLY the length the timeline
gives it (VideoCompileStepExecutor.SegmentLengthSuffix):
tpad=stop_mode=clone:stop_duration=<len>,trim=end=<len> — any shortfall is filled with the clip's
own last frame, and a segment that really has every frame is untouched. On the customer's two-clip
composite, frame 158 went from 16.5% raw green to none, and the whole four-clip composite has no
frame over 0.15%. SegmentLengthBoundaryTests reproduces it with real ffmpeg on a clip shaped like
the generated ones (and shows the unpadded graph starting the next clip a frame early).
What remains at the 0.15% level is a different thing: on the finger-swipe clip around 7.7 s, a hand crossing the phone's lower-right edge pulls the tracked edge inward for a few frames, leaving a thin strip of plate beside the hand — a tracking limit of the kind described in Tracking quality on the customer clips, not timing.
Inserting an existing video
An insert's content used to be only something the motion-graphics agent rendered in the same run —
the compile step forces such keys under the execution's own outputFiles/ prefix. There was no
way to say "put my clip on the phone", so a workflow asking for that had to list the clip as a
second VideoAnalyze source. That offered its shots to the story editor as footage, the editor
kept them, and the finished video was the phone take with the entire promo appended after it,
while the planner rendered a look-alike promo of its own for the screen.
Two paths now exist:
- Deterministic:
VideoCompileStepConfig.InsertContentProjectFileIdnames avideo/*orimage/*project file. WithEnableInserts, it is composited into every offered region that clearsMinInsertConfidence, most confident first, up toMaxInserts, skipping regions a plan already filled. No agent and noGraphicsPlanare needed — a tracking test isVideoAnalyze(the plate footage only,DetectInsertRegions: true) →VideoCompile. Still images are looped for the insert's whole window. - Planned: a
ScreenInsert'srenderedAssetStorageKeymay be a project file's id (or its exact storage key). It is matched against this project's own file list; anything else is treated as an agent render and re-anchored exactly as before. The planner's prompt now tells it to use the user's file instead of authoring a substitute, and the story editor's prompt tells it that a clip the request describes as content to place inside the picture is not footage to keep.
InsertFit (Auto | Stretch | Contain | Cover) shapes the content to the plate's measured
physical aspect before the warp maps its whole frame onto the plate. Auto always stretches an
agent render (it was authored for the plate, so it is unchanged) and a project file whose aspect is
within 15% of the plate's, and otherwise contains — a landscape promo plays letterboxed on a
portrait phone instead of squashed.
The EDL's appliedInserts records each insert's contentOrigin (plan/config),
contentProjectFileId and contentFit.
Different content per screen (InsertContents)
InsertContentProjectFileId puts ONE file on every plate, so a portrait UI for the phones and a
landscape UI for the laptop in the same edit took two compile passes. InsertContents (a list of
VideoInsertContentAssignment(ProjectFileId, SourceIndex?, RegionId?, Fit?)) assigns content per
plate on the deterministic path. For each offered screen plate (feature-tracked objects excluded, as
before), PickInsertContent takes the most specific matching entry — one naming the plate's
RegionId outranks one naming its SourceIndex (the analyze step's 0-based clip), which outranks
one naming neither; list order breaks ties — and falls back to InsertContentProjectFileId; a plate
neither covers gets no deterministic content. An entry's Fit overrides InsertFit for that
content. Every other rule (offered ids, confidence, MaxInserts, plan-filled plates first) is the
shared loop's, unchanged. Null/empty is byte-identical to the single-content behaviour.
Dips to black in insert content (InsertContentDips)
Screen recordings dip to black between screens (one measured at 2.8–3.6 s, 5.7–6.0 s, 8.6–9.1 s and
11.3 s), and on a tracked phone that reads as the phone switching off. For every project-file
video used as insert content (renders made for the plate and still images are never touched),
the compile runs ffmpeg's blackdetect once per file (d=0.08:pic_th=0.98:pix_th=0.10 — a frame
with 98% of its pixels under 10% luma, for at least 80 ms; parsed by ParseBlackDetect) and, per
InsertContentDips (InsertContentDipMode, append-only):
| Value | What is composited |
|---|---|
Freeze (default) | A copy where each dip's frames are dropped and fps=…:start_time=0 refills the gap by repeating the last good frame — the content keeps its length and timing; a dip at the very start shows the first good frame, one running to the end is held by tpad |
Cut | A copy with every dip removed (select + setpts) — the content gets shorter |
Keep | The file as it is; nothing is detected (no extra ffmpeg call) |
The copy is re-encoded once (libx264 -crf 16) and reused for every plate showing that file
(BuildInsertDipFilter). Content that is at least 90% black is left alone
(content_mostly_black); a failed detection or re-encode composites the original
(bridge_failed) — never a failed compile. Each applied insert whose content had dips carries, in
the EDL and output summary:
"contentDips": { "treatment": "Freeze", "intervals": [ { "startSec": 2.8, "endSec": 3.6 } ],
"totalSec": 0.8, "applied": true }
(absent when no dip was found, so an insert with clean content keeps its exact shape). Because the
default is on, a compile with insert content now makes one blackdetect pass per content file.
Tracking in the editor
The timeline manifest's insert items carry, besides the source-space quadKeyframes the
EditTimeline v2 seeder reads, outputQuadKeyframes — the corners the encode actually warped to,
in output timeline seconds and 0..1 of the output picture, after the cut, overscan and any
letterbox fit — plus matted and tuning. The program monitor draws those inside the video's
displayed rectangle. It previously multiplied source-space corners by the whole monitor canvas,
so a portrait video pillarboxed in the landscape monitor had its tracked corners drawn out in the
black bars, and it labelled any plate with non-zero measured curvature "mesh" even when it was
composited with a planar pin.
Tracking objects
VideoAnalyzeStepConfig.TrackObjects finds and follows arbitrary flat objects — a licence plate,
a sign, a screen with no green plate — with a markerless planar tracker
(Services/Video/Tracking/, pure C#, no new dependency). The chroma tracker above needs a
uniform-colour plate; this one needs only texture.
Pipeline, per source (ObjectTrackingService)
- Frames. One grayscale pass (
IFrameGridSampler.SampleGrayAsync, bicubic — the tracker needs gradients, which an area downscale blurs) atObjectTrackSampleFps(default 15), longest sideObjectTrackMaxSide(default 960). - Seeds. A target with a
Region(normalized corners atAtSec) seeds exactly there with no model call. A target with only aLabelis FOUND: on frames everyObjectDetectEverySec(default 1 s; pastMaxObjectDetectFramesthe lookups are spread evenly over the whole clip) theVisionprovider (IObjectLocator/VisionObjectLocator) returns boxes on a 0..1000 integer grid. Each box is then refined on a zoomed crop around it (CropAround: 2.5x the box, at least 12% of the frame, re-extracted at 768 px withExtractCroppedKeyframeAsync), because vision models place boxes loosely on a full frame and much more tightly when the object fills the image. The model contributes only seed boxes; the label is the author's text, sanitized to one short line before it enters the prompt. - De-duplication. Seeds are processed placed-first then in time order; a seed whose box overlaps (IoU > 0.3) the same label's existing track at that frame is the object that track already follows and is skipped. This is what turns "looked every second" into one track per object rather than one per lookup, while still catching objects that enter later or that a track lost.
- Tuning, once per label (
PlanarTrackerTuner): motion model (homography / affine / similarity) × LK window radius (5, 7, 10, 13). Each candidate tracks the seed up to 16 frames out, then tracks the quad it ended on back to the seed; score iscoverage × meanConfidence × exp(−roundTripError / 0.03)with the error as a fraction of the seed quad's diagonal. Pyramid depth was searched too at first and dropped: depths 3 and 4 gave bit-identical tracks for every candidate. - Tracking (
PlanarTracker), forward and backward from the seed, per frame:- Shi-Tomasi features in the quad grown by
DetectionMargin(15%), so the object's own outline corners count; background features caught in the margin move differently and RANSAC drops them. - Pyramidal Lucas-Kanade (Bouguet) frame to frame, keeping only points whose backward flow returns within 1 px (forward-backward check).
- RANSAC fit of the tuned model (deterministic seed, Hartley-normalized least squares).
- Drift control: the chained estimate is used only as the starting guess for registering the SEED frame's features directly onto the current frame; when that registration agrees, it replaces the chain, so error does not accumulate frame over frame. It falls back to the chain when the target has changed too much to register.
- Sanity (convex quad, per-frame area change within 35%, centre near the frame) and an appearance gate — NCC of the rectified quad against the seed and a slowly updated recent template. Below 0.3 for more than 3 frames ends the track instead of letting it slide onto the background.
- Features are replenished inside the quad when fewer than half survive, mapped back to seed space so they join the reference registration.
- Shi-Tomasi features in the quad grown by
- Output. Each run becomes a
VideoInsertRegionTrackwithMethod = "feature", itsLabel, 3-frame-smoothed keyframes,Confidence = coverage × mean per-frame confidence(per-frame: RANSAC inlier ratio × appearance), and theTuningrecord. Tracks join the source's insert regions after the chroma plates, continuing theirr{n}ids, so everything that consumes an insert region — screen inserts included — works on them. The view marks them"kind": "object"with theirlabel.
meta.objectTracking (applied, degraded, reason, tracks, visionCalls) appears only when
TrackObjects is configured. No Vision provider, a failed lookup, or a tracker exception all
degrade with a reason; none fails the step.
Measured
On a synthetic clip with exact ground truth — a licence-plate graphic composited onto the real
broll-04-servers.mp4 with ffmpeg's perspective filter, growing 35%, rotating and skewing over a
camera move in the other direction, 135 frames at 960x540 — the tuned tracker followed every frame
with a median corner error of 4.4 px at 1080p (p95 8.5, max 18). Untuned (homography, 21 px window)
it drifted to 49 px by the end, which is what the round-trip search exists to catch.
Censoring tracked objects
VideoCompileStepConfig.CensorLabels obscures every feature-tracked region whose label matches
(case-insensitive; "*" = all), CensorStyle Blur | Pixelate | Fill.
- Every portion.
OutputTimeline.MapWindowToOutputAllmaps a track into EVERY output stretch its source window survives into. Inserts use the first stretch only; for privacy that would leave an object visible after a cut. - No confidence gate. A doubtful track is still obscured — over-obscuring is the safe failure.
- Padding and hold. Each region gets a margin of
CensorPadding(default 0.25) times its LONGER side on every side — a licence plate is ~4x wider than tall, and a margin proportional to its own height added only a few pixels where a vision box a little off vertically needed the most cover (on the car-meetup clip a box 16 px too high left the bottom row of characters readable). Each track's first/last position is heldCensorHoldSec(default 0.4 s) outward. - Rendering (
CensorFilterBuilder): the picture is split and the copy obscured ONCE (Gaussian blur sigma 1.5% of the canvas width; 2.5%-of-width mosaic; or solid black). One combined mask is drawn on a quarter-resolution black canvas: per region and per one-second chunk of its track, a white box as large as the region's largest padded extent in that chunk, moved every frame by anoverlayx/y expression in output time (a balancedif(lt(t,…))tree, since ffmpeg rejects deeply nested expressions). The mask is scaled up, alpha-merged with the obscured copy and overlaid back. Twelve plates on 16 s of 1080p: 64 s when every region obscured the whole frame, 9 s now. - Why boxes, not the insert warp. The first version pinned a white plate onto each region
with
perspective, exactly like a screen insert. ffmpeg'sperspectiveprecomputes fixed-point coordinate maps that overflow when a full-frame plate is squeezed ~20x onto a plate-sized quad; the "mask" covered everything above-left of each plate. Screen inserts never hit this because screens are large. Axis-aligned boxes over the padded quad need no warp, and erring large is the right bias for privacy. - Stage order. It is the last picture stage before the program fade, so anything composited onto the object is obscured too; on the segmented path it runs on the program body, before generated cold opens/end cards are concatenated, so its body-timeline windows stay aligned.
- Fail closed. A compile with
CensorLabelsfails withCENSOR_UNVERIFIEDinstead of producing a video that may show what it was meant to hide when: the analyze step tracked no objects at all; any censor label (other than*) is not one the analyze step looked for in every source (Provenance.ObjectTrackingLabels, compared case-insensitively — "licence plate" vs "license plate" is a mismatch, deliberately); or object tracking DEGRADED — no Vision provider, a failed or unreadable lookup (a refusal or truncated reply is a failure, never "nothing here"), more than 16 instances in one lookup, a tracker error, or sightings left untracked becauseMaxObjectTrackswas reached. A clean run that genuinely found nothing compiles, and the EDL'scensor.unmatchedLabels(also reported as a progress message) names each label that matched nothing.CensorLabelswithMode = StreamCopyis a hard failure (CENSOR_REQUIRES_REENCODE). Outputs whoseobjectTracking,visionorinsertTrackingstage degraded are never written to the step cache (StepResultCache.ReportsDegradation), so a fixed provider is not masked by a cached failure. Transcription is deliberately excluded: it reports "degraded" for a source with no audio or an install with no speech-to-text provider, which are permanent facts of the input. - The timeline manifest carries a
censortrack whose items haveoutputQuadKeyframes(the tracked quad, unpadded); the editor draws them as red dashed outlines with label and confidence.
Measured on real footage
vintage-cars-meetup.mp4 (Wikimedia Commons, "Renkontiĝo de ŝatantoj de malnovaj aŭtoj en Tjumeno
(2022)", CC BY-SA 4.0): 16 s handheld, two parked cars, people walking in front of both plates.
36 vision lookups (DeepSeek Vision, via the plain-JSON fallback), 17 tracks; both plates are
obscured in every sampled frame from 0 to 16 s. Over-detections (a turn signal, a patch of
pavement) are obscured too; that is the intended bias.
KeepWholeSources
VideoCompileStepConfig.KeepWholeSources ignores Decision and keeps every analyzed source whole,
in source order (WholeSourcesDecision, built from the artifact's own shot ids and checked against
every shot rather than the offered set). It is the no-agent path for compiles that add to footage
rather than cut it: a tracking test, a censor pass, a screen insert of a project file.
The edit room
StepType.EditRoom replaces the single AgentType.VideoStoryEditor decision step with a
multi-agent deliberation: several editor-role seats plus a director converse in a live
Microsoft.Agents.AI.Workflows group chat over the same bounded VideoAnalyze view a solo editor
would see, and the director synthesizes their discussion into ONE schema-validated
VideoEditDecisionOutput — the exact same schema and the exact same rushcut invariant (never a
timestamp, only offered ids) AgentType.VideoStoryEditor already produces. VideoCompileStepExecutor
needs zero changes to consume it: Decision just points at the EditRoom step's StepOrder
instead of a solo VideoStoryEditor step's.
┌─────────────────┐ ┌──────────────────────────┐ ┌──────────────────┐
│ StepType. │ │ StepType.EditRoom │ │ StepType. │
│ VideoAnalyze │────▶│ several editor seats + │────▶│ VideoCompile │
│ │ │ a director, live group │ │ │
│ (unchanged) │ │ chat, then one synthesis │ │ (unchanged) │
│ │ │ call outside the chat │ │ │
└─────────────────┘ └──────────────────────────┘ └──────────────────┘
emits VideoEditDecisionOutput
(+ an additive "room" metadata
block) — identical shape to a
solo VideoStoryEditor step
Why a group chat, and why it's deterministic-scheduled
Every seat sees the identical bounded view and can speak to any part of it — there is no
information asymmetry between seats that would make an LLM-driven "who should speak next" routing
decision meaningful. EditRoomGroupChatManager (WorkflowEngine/Agents/EditRoom/EditRoomGroupChatManager.cs
— since the graphics room landed, a thin binding of the room-generic RoomGroupChatManager base,
contributing only the edit room's [sgt]{n} offered-id regex; see
The shared room infrastructure)
is therefore a plain, deterministic round-robin scheduler over the configured seats, followed by
the director: it makes zero model/network calls itself, only orchestrating which
already-constructed AIAgent speaks next. Routing this through an LLM would double the room's cost
for a decision that doesn't need intelligence.
The two agents
- The editor seats (
EditRoomStepConfig.Seats, three by default —PacingEditor/StoryEditor/CraftEditor, argued personas for rhythm/narrative/material-quality respectively) all resolve to the SAME built-inAgentType.VideoStoryEditoragent (its own seeded prompt/tools/provider) — no newAgentTypeenum member exists per seat, since only the persona differs, and personas are injected per-turn, not baked into separate agent definitions. AgentType.VideoEditDirectoris used TWICE, through two different code paths: once per-turn as a ROOM PARTICIPANT (free-form prose, moderates disagreement, ends the room by emitting the literal sentinelROOM_DECIDEDonce satisfied), and once more, entirely OUTSIDE the group chat, for a single ordinary structured-output SYNTHESIS call (ReelBoltAgentBase.RunAsync, the same mechanism every other structured-output agent uses) that converts the room's transcript into the finalVideoEditDecisionOutput. Same minimal read-only tool scope asVideoStoryEditor(no sandbox, no write/render tools — it only decides).
How a turn is built: RoomSeatAgent
RoomSeatAgent (WorkflowEngine/Agents/Rooms/RoomSeatAgent.cs — named EditRoomSeatAgent until
the graphics room landed; renamed unchanged since it was already fully room-agnostic) wraps each
already-constructed inner AIAgent (built by EditRoomStepExecutor via the same chat-client-
resolution path ReelBoltAgentBase.CreateAgentAsync uses) as a DelegatingAIAgent, overriding
BOTH RunCoreAsync and RunCoreStreamingAsync — the group chat host always invokes participants
through the STREAMING path, so a wrapper that only overrides the non-streaming one is silently
bypassed (measured live against the real rc2 package). Three things happen on every turn:
- Per-turn sampling options are injected, since the group chat host always passes
options == nullto a participant — this wrapper is the only way to control temperature/reasoning-effort/MaxOutputTokensper turn (EditRoomStepConfig.Temperature/DirectorTemperature/ReasoningEffort/MaxTurnTokens), via the sameRawRepresentationFactoryOPENAI001mechanismReelBoltAgentBase.BuildChatOptionsalready uses forreasoning_effort.
- The seat's persona is appended as the LAST message, after the identical
[system instructions][bounded-view opening message]prefix every seat/the director share — measured live as a ~2-3x latency win, since it lets that identical prefix keep hitting the backend's prompt-prefix cache. Putting per-seat identity earlier defeats the cache; this ordering must not be changed. - A seat's turn throwing never aborts the room. Wrapped in try/catch: on failure, a synthetic
"[{SeatName} had no input this round]"turn is recorded instead, and the room continues.
Scheduling and termination
EditRoomGroupChatManager.SelectNextAgentAsync round-robins the editor seats for
EditRoomStepConfig.Rounds full passes, then always returns the director:
editorCount = seats.Count
turn i: i < editorCount * Rounds → editors[i % editorCount]
otherwise → director
EditRoomStepConfig.Termination gates how the room can end EARLY, on top of the hard
MaxTurns ceiling (mapped to GroupChatManager.MaximumIterationCount, clamped 2..20 — deliberately
far below the framework's own default of 40, which is a multi-hour runaway on this backend, not a
safety net):
| Mode | Behavior |
|---|---|
SentinelOnly | Ends when the director's most recent turn contains the literal ROOM_DECIDED token. |
Converged | Ends when the offered-id-vocabulary mentions across the last MinConvergenceRounds consecutive editor rounds are identical — the seats have stopped proposing anything new. |
SentinelOrConverged (default) | Either check ends the room early. |
FixedTurns | Neither check runs — only the MaxTurns ceiling ends the room. |
Both checks are evaluated in ShouldTerminateAsync, called BEFORE any agent has spoken too
(iteration 0, history = just the opening message) — the checks must not assume at least one turn
has happened, and don't. UpdateHistoryAsync is a pure pass-through: it observes every new message
for the executor's own transcript/progress bookkeeping but always returns the input history
unchanged — returning anything else would corrupt the framework's canonical transcript, since this
return value is NOT a per-turn view.
The executor: EditRoomStepExecutor
Same never-throws, always-valid-JSON discipline as VideoAnalyzeStepExecutor/
VideoCompileStepExecutor (see The three-stage shape):
- Deserializes
EditRoomConfigJson; resolves the bounded view viaEditRoomStepConfig.View(anExtractInputRef, reused verbatim fromVideoAnalyzeStepConfig— the samePrevious/StepresolutionVideoCompileStepExecutor.ResolveDecisionJsonalready established forDecision/GraphicsPlan/MusicPlan) and extracts the offered shot/silence/segment id vocabulary from it. - Builds every seat's and the director's
AIAgent, wraps each inRoomSeatAgent, and runs them throughAgentWorkflowBuilder.CreateGroupChatBuilderWith(...).AddParticipants(...).Build()viaInProcessExecution.RunStreamingAsync, bounded byEditRoomStepConfig.RoomTimeoutSeconds. The bounded view is sent as ONE opening chat message (not folded into the agent instructions), since that message becomes the shared prefix every seat's every turn hits the prompt-prefix cache against. - Reads the transcript back off the
WorkflowOutputEventthe workflow emits once it completes. - Makes ONE standalone structured-output synthesis call (
AgentType.VideoEditDirector.RunAsync, outside the group chat) with the view + rendered transcript, retried up toMaxSynthesisAttemptson an empty/unparseable result. - Validates deterministically, never trusting the model: any
Keepspan whoseFromId/ToIdisn't in the offered-id set is dropped, and — when the view spans more than one source clip (see Multiple source clips) — so is any span whoseFromId/ToIdcarry two different"src"indices, the exact shapeVideoCompileStepExecutorrejects hard asMIXED_SOURCE_SPAN(dropping it here instead keeps the step's degrade-not-fail discipline, and the remaining single-clip spans still compile). Both are recorded in the output'sroom.droppedSpanCount; the mixed-source subset also inroom.droppedMixedSourceSpanCount. A synthesized decision containing a mixed-source span first gets a retry with precise feedback (withinMaxSynthesisAttempts) before the drop is accepted. The room charter prompt every seat (and the director's room-participant turns) runs under states the same one-hard-rule the soloVideoStoryEditorprompt's "Multiple source clips" section does — a single kept run's first and last id must come from the SAME clip — guarded byRoom_charter_prompt_carries_the_multi_source_single_clip_span_rule. - If the room failed outright, produced zero usable turns, the synthesis call never produced a
usable decision, or every Keep span got dropped as unoffered — and
FallbackToSoloEditor(defaulttrue) — falls back to ONE ordinary soloAgentType.VideoStoryEditorcall: today's existing single-editor pipeline, unchanged. The step still completes successfully with a valid decision;room.degraded/room.degradeReasonrecord what happened. - Emits
output_json= the validatedVideoEditDecisionOutput, camelCase-serialized, plus an ADDITIVE sibling"room"object (seats, rounds, turn count, how the room terminated, whether it converged, synthesis attempts, dropped-span count, whether it degraded and why, the transcript artifact's storage key). This is safe becauseVideoCompileStepExecutor's own deserialization ofVideoEditDecisionOutputuses noUnmappedMemberHandling.Disallow— the extra"room"key is silently ignored by that step exactly like any other consumer expecting the plain schema.
The transcript artifact — NEVER authoritative
When EditRoomStepConfig.PersistTranscript (default true), the full, unabridged room transcript
is uploaded as this step's ArtifactStorageKey, under the same projects/{projectId}/agentFiles/ video-analysis/{executionId}/step-{order}-room-transcript.json prefix convention VideoAnalyze/
VideoCompile already use for non-playable JSON artifacts (see Where artifacts
live) — reachable through the same GET /api/v1/projects/{projectId}/ step-results/{stepResultId}/artifact endpoint. This transcript is free-form model prose and must
never be parsed back into a decision. A seat or the director can say something that sounds like a
timestamp in passing prose ("that pause feels like it's about three seconds") — harmless as
commentary, but nothing downstream may ever try to extract a number from it. The only authoritative
output of an EditRoom step is the synthesized, deterministically-validated
VideoEditDecisionOutput on output_json.
Live progress and the transcript event
Every completed turn reports a context.ReportProgressAsync stage/percent update (e.g.
"PacingEditor is proposing a cut (turn 3/8)" / "Director is reviewing") — the same ephemeral,
supersedable WorkflowStepProgress signal every other long-running step uses. Additionally, when
EditRoomStepConfig.StreamTurns (default true), each completed turn also publishes an
append-only WorkflowStepChatTurn integration event — a structural twin of
WorkflowStepReasoningCaptured (sequence-numbered, every turn preserved), deliberately NOT a reuse
of WorkflowStepProgress, whose "a later progress event supersedes any earlier one" semantics are
wrong for a transcript where every turn matters. WorkflowStepChatTurn.IdsMentioned is
server-extracted (the same regex EditRoomGroupChatManager uses for convergence checking, filtered
against the offered-id set) — display/audit only, never trusted as the decision itself. The Go
API/frontend relay of this new event type is a separate, follow-up pass.
Config reference
| Field | Default | Notes |
|---|---|---|
Version | 1 | Config schema version |
View | null | ExtractInputRef (Previous/Step only, reused verbatim from VideoAnalyzeStepConfig) — which step's bounded {view, meta} envelope every seat and the director see. null resolves to Previous |
Seats | null | IReadOnlyList<EditRoomSeat> — null/empty resolves to the three validated-live defaults (PacingEditor/StoryEditor/CraftEditor), each {Name, Persona, AgentDefinitionId} |
Rounds | 2 | Full round-robin passes over every seat before the director speaks |
MaxTurns | 8 | Hard ceiling (GroupChatManager.MaximumIterationCount), clamped 2..20 — NOT the framework's own default of 40 |
DirectorAgentDefinitionId | null | Per-agent-definition override for the director seat |
Termination | SentinelOrConverged | SentinelOnly / Converged / SentinelOrConverged / FixedTurns — see the table above |
MinConvergenceRounds | 2 | Consecutive rounds with an unchanged offered-id-mention set required for Converged/SentinelOrConverged to fire |
MaxTurnTokens | 220 | Per-turn MaxOutputTokens, injected via RoomSeatAgent |
MaxHistoryChars | 40000 | Soft cap on how much of the room transcript is rendered into the synthesis prompt (oldest turns dropped first beyond this) |
Temperature | 0.7 | Sampling temperature for every editor seat's turn |
DirectorTemperature | 0.3 | Sampling temperature for the director's ROOM-PARTICIPANT turns only — the standalone synthesis call uses VideoEditDirectorAgent's own AgentModelSettings default (0.3/"low") instead |
ReasoningEffort | "none" | Injected into every seat's (and the room-participant director's) turn — measured ~3x latency reduction on this backend. Must pass ReelBoltAgentBase.ValidReasoningEfforts |
RoomTimeoutSeconds | 1200 | Hard wall-clock budget for the whole group-chat run (not the later synthesis call) |
PersistTranscript | true | Whether the full transcript is uploaded as ArtifactStorageKey — see above |
StreamTurns | true | Whether a WorkflowStepChatTurn event is published per turn (progress reporting always happens regardless) |
FallbackToSoloEditor | true | Whether a failed/empty room decision falls back to one ordinary solo AgentType.VideoStoryEditor call |
MaxSynthesisAttempts | 2 | Retry attempts for the standalone structured-output synthesis call |
The video-derush-edit-room template
A sixth opt-in template (AutoCreateOnProject: false), replacing video-derush-edit's middle
Agent(VideoStoryEditor) step with the room: VideoAnalyze (Source: ProjectFile) → EditRoom
(View: Previous) → VideoCompile (Decision: Step 2, AnalysisStepOrder: 1) → ReviewLoop
(AgentType.VideoReviewAgent, looping back to step 2). The EditRoom step's own AgentDefinitionId
FK is satisfied by the same AgentType.VideoTransform deterministic placeholder VideoAnalyze/
VideoCompile steps already use — the room's real seats/director are resolved independently, from
EditRoomConfigJson, never from the step's own AgentDefinitionId. Deserialization-tested the same
way as the other video templates (WorkflowTemplateCatalogConfigDeserializationTests.cs).
The shared room infrastructure
The edit room's mechanics were extracted into a room-generic base the moment a second room (the graphics room, below) needed them — deliberately as ONE shared implementation, not per-room copies, since multi-agent deliberation is intended as a standing pattern in this codebase (a color-grading room is the next planned consumer). The seams:
RoomSeatAgent(WorkflowEngine/Agents/Rooms/RoomSeatAgent.cs, formerlyEditRoomSeatAgent— renamed unchanged, it was already fully room-agnostic) + itsRoomTurnResultrecord: theDelegatingAIAgentwrapper providing per-turn sampling-option injection, the prefix-cache-preserving persona-last message ordering, seat display names, and turn-failure containment. A new room reuses it as-is.RoomGroupChatManager(WorkflowEngine/Agents/Rooms/RoomGroupChatManager.cs): the deterministic scheduler/terminator base — round-robin over N rounds then the director, theROOM_DECIDEDsentinel check, offered-id-mention convergence, the explicit ceiling check, and turn observation. A concrete room contributes ONLY its offered-id vocabulary: a thin subclass (EditRoomGroupChatManagerbinds[sgt]\d+,GraphicsRoomGroupChatManagerbindsp\d+) passing its compiled regex to the base constructor, plus a staticExtractOfferedIdMentionsconvenience bound to that regex.IRoomStepConfig(Shared/Workflows/RoomStepConfig.cs): the config surface the shared infrastructure reads (view ref, effective seats, rounds/turn ceiling, termination knobs, temperatures, reasoning effort, timeout, transcript/stream flags,FallbackToSolo, synthesis attempts). Each room's own JSON record (EditRoomStepConfig,GraphicsRoomStepConfig) implements it on top of its unchanged JSON shape —EditRoomSeatandEditRoomTerminationModeare the room-GENERIC seat record and termination enum despite their names, kept under their original names so no persisted config or test broke when the base was extracted.RoomStepExecutorBase<TDecision>(WorkflowEngine/Execution/StepExecutors/RoomStepExecutorBase.cs): the executor template — never-throws/always-valid-JSON discipline, config/view resolution, the group-chat run (agent construction, progress +WorkflowStepChatTurnevents with cumulative token tallies, theTurnTokenkickoff, the room timeout), transcript persistence (both the object-store artifact and the DB-persistedChatTranscriptJson), the retried standalone synthesis call, the solo-agent fallback, offered-id filtering, and the additive"room"metadata block onoutput_json. A concrete room supplies: itsStepType/config column/config type, offered-id extraction + mention regex, charter prompt + director turn directive, the seat/director/soloAgentTypes, decision normalize/reject/filter hooks (the edit room rejects an emptyKeeplist; the graphics room accepts an empty plan), and optional overrides for progress-label wording, room-turn tool scope (GetRoomTurnTools), view enrichment (PrepareViewAsync), and the empty-view outcome (BuildDecisionForEmptyView).- The workflow's user request reaches all three room prompts. A room does not go through
StepExecutionContext.BuildAgentInput, so it does not get the--- User Request:block an ordinaryStepType.Agentstep's prompt ends with for free; for a while it got it nowhere, and anEditRoomstep — documented as a drop-in replacement for a soloAgent(VideoStoryEditor)step — silently discarded the only channel a workflow has for creative direction. Every seeded room template then ran withRequiresUserInput: false(every template now asks for the brief), which is why nothing surfaced it until a real run asked the graphics room for two tracked inserts and the colour-grade room for a violet grade, and got zero inserts andLook: "None"with every step green. The opening message, the synthesis prompt and the solo-fallback prompt all now go throughStepExecutionContext.WithUserRequest, sharing one wording with the agent-step path; the opening message is composed at the CALL SITE so a room overridingBuildOpeningMessagecannot drop it again. With no user request every prompt is byte-identical to before.
To build a new room (e.g. color grading): add a StepType + {X}RoomConfigJson jsonb column
(mapped in BOTH DbContexts, migration on the WorkflowEngine context only — it owns
workflow_steps), a config record implementing IRoomStepConfig, a RoomGroupChatManager
subclass binding the room's id regex, a RoomStepExecutorBase<TDecision> subclass binding the
hooks above, a dual-role director agent (+ seeded prompt with the verbatim-consistency test), and
a template. The synthesis output schema should be an EXISTING single-agent schema whenever a solo
equivalent exists, so downstream consumers need zero changes — that is the entire trick that let
VideoCompileStepExecutor consume both rooms' outputs untouched.
The room convening gate — ColorGrade room only
A room can hand a System One decision model a set of questions instead of (or alongside) just
convening the group chat. The surface is IRoomStepConfig.ConveneGate (Off / Shadow / Gate,
default Off everywhere) plus three virtual seams on RoomStepExecutorBase<TDecision>:
BuildConveneQuestions (the wire questions a pre-room gate would ask over the bounded view),
BuildConveneAgenda (the text an escalating gate would append to the opening message), and
BuildShadowQuestions (the post-synthesis questions the Shadow arm records, each carrying the
room's own answer as SystemTwoAnswer). All three default to null, and null is what "this room
has no gate" means.
It ships for the ColorGrade room only. That is a deliberate boundary, not a partial implementation:
- A System One
choice/score/noulquestion can only produce a decision that decomposes into a fixed option set.ColorGradePlanOutputis exactly that — one whole-program grade in enum words — soColorGradeRoomStepExecutorsupplies the four questionsAgentType.Coloristalready poses (Colorist.Look/Colorist.Strength/Colorist.ShadowTone/Colorist.HighlightTone), copied verbatim fromAgentStepExecutor.BuildColoristQuestions. The site strings are the solo agent's on purpose: the calibration view keys its entries by(Site, Question), so a room step's observations and a solo colorist step's land on ONE series and aggregate together. Only the recordedAgentTypediffers — the room's ownColorGradeDirector— which keeps a room observation attributable to the room rather than passing it off as a solo one. EditRoomStepExecutorandGraphicsRoomStepExecutoroverride all three seams back tonull, with the reason stated in the code next to the override. Neither decision decomposes into fixed options:VideoEditDecisionOutputis an id-anchored keep/drop list over whatever ids this video's analyze step happened to offer, andMotionGraphicsPlanOutputis an open, generative list of overlays whose entries carry free text and may carry a self-rendered asset key. A question set for either could only gate on a proxy for the decision, and gating on a proxy is worse than not gating: it would trade the room's real decision for a merely correlated one while looking like the room had been gated. The overrides are explicit so the absence reads as a decision rather than an oversight.- Their
ConveneGateconfig member still exists and still defaults toOff, so setting it toGateon an edit or graphics room is a no-op rather than an error — the room convenes exactly as it always has.RoomConveneGateTestsasserts the emitted output is byte-identical to the same step withOff, using a strict gate mock that fails on any invocation at all. Gateis live for the grade room. WithConveneGate = Gateand aCapability = Decisionprovider configured, the arm resolves at the dispatch point inRoomStepExecutorBase<TDecision>.ExecuteAsync— after the offered-ids gate andRecordResolvedInput, beforeRunRoomAsync— by callingIDecisionGate.DecideAsyncwith aDecisionGateRequestwhoseStateis the bounded view alone (not the workflow's user request, not the prior-step output history), carryingSystemTwoAnswer: nulland noAgentTokensUsedbecause nothing has run yet. It no-ops to "convene" unless the mode isGate, anIDecisionGatewas injected, and the room supplies a non-emptyBuildConveneQuestions— so a null gate stays the legal, inert configuration it has always been. The gate gets its owntry/catch: onlyOperationCanceledExceptionon a genuinely cancelled context is rethrown, and any other failure logs and convenes.- A fully accepted outcome skips the room.
DecisionGateOutcome.Acceptedis all-or-nothing — every question cleared bothAcceptAtandMinMargin— and means the group chat does not run. The step takes the room's existingRunSoloFallbackAsync, reused unchanged rather than reimplemented, so a non-convened step produces the same decision shape, the same offered-ids filtering and the same room-specific validation a degraded one does; and it advances exactly as a convened step does, through the sameSuccesscall, withroom.terminationReason = "not-convened",turnCount 0and no transcript. That is not a degrade — skipping the room is the gate working as designed — sodegradedstaysfalseanddegradeReasonstays null. A solo fallback that itself fails is the room's existing{FailureCodePrefix}_FAILEDfailure mode, never a silently emitted decision. - Everything else convenes the room, exactly as before. An escalation, no decision at all
(
outcome.HasDecision == false— no Decision provider row resolved, the gate-level deadline arriving asDecisionGateOutcome.NoDecisionrather than throwing, any failure), a nullIDecisionGate, and an empty question set all fall through to the ordinary group chat. On escalation the resulting agenda is appended to the opening message at the CALL SITE, afterWithUserRequestandWithRetryGuidance, so a room overridingBuildOpeningMessagecannot drop it and a retried room is handed the same agenda again; null/whitespace is a strict no-op that leaves the message byte-for-byte what it was. - The grade room supplies the agenda through the shared formatter.
ColorGradeRoomStepExecutoroverridesBuildConveneAgendato call the base'sRenderConveneAgenda(BuildConveneQuestions(viewRoot), answers)— the question set is room-specific, the wording is not. Per split question, most-contested-first, it emitsThe fast model split on Look: Filmic 0.44 vs Warm 0.41. Resolve that disagreement first., culture-invariant and bounded byMaxConveneAgendaChars(600). That cap is a small fraction of the room's 40 000-characterMaxHistoryCharsdeliberately: the agenda is turn zero of the group chat, so every character of it is re-sent on every turn and is rendered into the synthesis prompt alongside the bounded view under that same transcript budget — a wide question set must be able to lose its tail rather than crowd the room's own history out of it. It keeps whole sentences and then stops (never truncating an option name or a probability) and returns null when nothing qualifies, so an all-confident, empty or degenerate answer set leaves the opening message byte-for-byte what it was. - The
Shadowarm is live too: withConveneGate = Shadowand aCapability = Decisionprovider configured, the grade room records one observation per question after synthesis and changes nothing about its decision, status or transcript. No shipped configuration enables either mode — every example is commented out andOff.Offis the default everywhere and is byte-identically today's behaviour, and a room runs with no Decision provider configured at all: that is the normal local case, not an error. The defaultConveneAcceptAt/ConveneMinMarginvalues (0.85/0.25) are phase-1 fallbacks, not tuned or calibrated thresholds — there is noCapability = Decisionprovider row and zerodecision_observationsrows in this environment to calibrate them against.
The graphics room
StepType.GraphicsRoom replaces the single AgentType.MotionGraphicsPlanner planning step with a
multi-agent deliberation over the SAME offered candidates a solo planner consumes — both the
view.placements overlay candidates (p{n} ids, Phase 3) and the view.insertRegions tracked
chroma plates (r{n} ids, Phase 5): several motion-graphics-artist seats plus a lead-artist director
(AgentType.MotionGraphicsDirector) converse in a live group chat, then the director synthesizes
ONE schema-validated MotionGraphicsPlanOutput — the exact schema, and the exact extended rushcut
invariant (never a timestamp OR a pixel coordinate, only offered placement ids, reflection-tested
by MotionGraphicsPlanOutputInvariantTests unchanged), the solo planner already produces.
VideoCompileStepExecutor needs zero changes: GraphicsPlan just points at the
GraphicsRoom step's StepOrder, and its deserialization skips the additive "room" metadata
key exactly as Decision resolution does for the edit room.
Everything structural is the shared room infrastructure; what is specific to this room:
- The seats (
GraphicsRoomStepConfig.DefaultSeats, four by default —LayoutArtist/TimingArtist/CopyArtist/InsertArtist) divide the actual decision space: WHERE (regions, fit scores, light/dark hints, clutter), WHEN (which candidate window lines up with what is said or shown,inEditsurvival, duration words), WHAT/HOW (copy brevity, rendered-graphic vs plain text, and whether the right number of overlays is zero), and WHICH PLATES (which offeredinsertRegionsare real screens rather than low-confidence false positives, what each should show given itsaspectandsurfaceword, and when leaving a plate empty is right). All four resolve to the same built-inAgentType.MotionGraphicsPlanneragent — personas are injected per-turn, exactly the edit room's one-agent-many-personas pattern. - The fourth seat is not decoration. The room originally had three, all of them arguing
placements, and every other room-specific string — charter, director directive, synthesis
instruction,
ExtractOfferedIds,FilterToOfferedIds— was overlay-only too. Run against footage carrying two real tracked plates, with a user request explicitly asking for an insert on each, the room discussedp{n}candidates for all eight turns and synthesized an emptyinsertslist: two raw green rectangles in the finished film, with the room, the compile step and the execution all reporting success. Nothing owned that half ofMotionGraphicsPlanOutput, so nothing raised it. The charter now describes inserts as the room's second decision, the director is told that silence on offered plates is not convergence, and the synthesis instruction names both lists. - Both id vocabularies, filtered separately.
ExtractOfferedIdsreturnsp{n}andr{n}together — so convergence detection covers the whole decision and a plates-only view (graphics placements off, insert tracking on) still puts the room to work rather than short-circuiting to an empty plan.FilterToOfferedIdsthen re-reads the two lists from the view and filters each kind against its OWN, because the combined set would wave through an overlay anchored tor0or an insert pointing atp3. AgentType.MotionGraphicsDirectoris used TWICE, likeVideoEditDirector: as the room-participant moderator (prose,ROOM_DECIDED), and for the standaloneMotionGraphicsPlanOutputsynthesis call. UNLIKEVideoEditDirector, its grant is the full sandbox+Remotion+render set minusWriteProjectFile(identical toMotionGraphicsPlanner's) — capability parity, so a room-planned overlay can still be backed by a real rendered transparent asset (RenderedAssetStorageKey) in the synthesis role. The prompt-injection tradeoff documented forMotionGraphicsPlanner(media-derived view text reaching a code-executing agent, bounded by the sandbox's containment) applies identically and is accepted for the same reasons.- Room turns are tool-restricted.
GraphicsRoomStepExecutor.GetRoomTurnToolsfilters every in-room agent (seats AND the director's room instance) to theProjectRead+WorkflowControlsubset of its grant (names taken fromToolGroupCatalog, so the subset cannot drift) — a ~220-token prose turn must never reach the sandbox; only the standalone synthesis call can. - The view is the analyze step's placements envelope, enriched with
inEdit. The template pointsViewexplicitly at the analyze step (Step 1—Previouswould resolve to the story editor's decision), andPrepareViewAsyncruns the sameIMotionGraphicsPlacementAnnotatora solo planner's prompt gets, read back throughStepExecutionContext.OutputForPrompt. A failed annotation degrades to the plain view, never fails the step. - An empty plan is a VALID outcome, twice over. A view offering zero placement ids AND zero
insert region ids completes
immediately with an empty plan (
room.terminationReason = "empty-view") instead of failing VIEW_UNRESOLVED, and a synthesized/solo plan with zero overlays (or filtered to zero by the offered-id check) completes normally — "prefer zero overlays over a cluttered edit" is the planning contract, so zero must never be treated as failure. The edit room's opposite choice (an emptyKeeplist degrades/fails) is the single biggest behavioral difference between the two rooms' validation hooks. - Failure codes:
GRAPHICS_ROOM_CONFIG_INVALID/VIEW_UNRESOLVED/GRAPHICS_ROOM_FAILED/UNEXPECTED_ERROR, mirroring the edit room's. Solo fallback is one ordinaryAgentType.MotionGraphicsPlannercall (FallbackToSoloPlanner, defaulttrue).
GraphicsRoomStepConfig (graphics_room_config_json, jsonb on BOTH DbContexts) carries the same
knobs as EditRoomStepConfig (see Config reference) with
FallbackToSoloPlanner in place of FallbackToSoloEditor, and one different default: MaxTurns
is 10, not 8. The scheduler gives the seats Seats * Rounds turns and the director every turn
after that, so with four seats and two rounds a ceiling of eight is consumed entirely by seat
turns — the director would never speak, never moderate, and never emit ROOM_DECIDED.
The video-derush-edit-graphics-room template
A seventh opt-in template (AutoCreateOnProject: false), replacing video-derush-edit-graphics's
Agent(MotionGraphicsPlanner) step with the room: VideoAnalyze (Source: ProjectFile,
emitOverlayPlacements: true) → Agent(VideoStoryEditor) → GraphicsRoom
(View: Step 1) → VideoCompile (Decision: Step 2, AnalysisStepOrder: 1,
enableGraphics: true, graphicsPlan: Step 3) → ReviewLoop(VideoReviewAgent) looping back to
step 2. The GraphicsRoom step's own AgentDefinitionId FK is satisfied by the same
AgentType.VideoTransform placeholder the other deterministic-config step types use.
Deserialization-tested in WorkflowTemplateCatalogConfigDeserializationTests.cs.
Color grading
Optional whole-program colour grading for the compiled edit, decided in enum WORDS only and resolved to concrete ffmpeg filter parameters entirely by first-party code. Two halves:
- The compile integration (
VideoCompileStepConfig.EnableColorGrade/ColorGradePlan+ColorGradeFilterBuilder) — deterministic, opt-in, soft-failure throughout, following theEnableGraphics/GraphicsPlanandEnableMusic/MusicPlanprecedent exactly. - The deciders — a solo
AgentType.ColoristAgent step, or the multi-agentStepType.ColorGradeRoom(below); both emit the exact sameColorGradePlanOutputshape, so the compile step consumes either interchangeably.
ColorGradePlanOutput: the words-only decision
The rushcut invariant, extended a fourth time (after cut ids, overlay placements, and music):
ColorGradePlanOutput is six plain strings — Look, Strength, ShadowTone, HighlightTone,
Reason, PlanRationale — and nothing else, pinned by ColorGradePlanOutputInvariantTests
(no numeric/time-bearing property; every property a string; the exact property set pinned so a
future id-bearing addition fails the suite). A colorist agent is structurally incapable of
emitting an RGB value, a curve point, a gamma/gain/contrast number, a percentage, or a
timestamp — its entire contribution is:
| Word | Values | Resolved to |
|---|---|---|
Look | None | Warm | Cool | Filmic | Vibrant | Muted | Mono | A named first-party colorbalance/eq/hue parameter set. None declines the whole grade — the compiled bytes then carry no grade filter whatsoever, tone words included |
Strength | Subtle | Normal | Strong | A 0.5/1.0/1.5 multiplier over the look's parameter DELTAS (Mono's hue=s=0 desaturation deliberately never scales) |
ShadowTone | Neutral | Lifted | Deepened | A colorlevels black-point input shift |
HighlightTone | Neutral | Softened | Brightened | A colorlevels white-point output/input shift |
Unknown Strength/ShadowTone/HighlightTone words normalize to Normal/Neutral/Neutral
(the same silent-normalize rule music applies to Intensity/Ducking); an unknown Look is
never guessed at — it degrades the whole grade as unknown_look_word, because the contract's
value is that every applied number traces to a recognized word.
ColorGradeFilterBuilder: the deterministic mapping
Pure static C# (Services/Video/ColorGradeFilterBuilder.cs), exact-string-tested
(ColorGradeFilterBuilderTests). Only four battle-tested core filters are ever emitted —
colorbalance (colour cast), eq (contrast/saturation/gamma), hue=s=0 (Mono), colorlevels
(black/white-point shaping, both tone words folded into ONE instance) — composed in that fixed
order. All numbers route through FfmpegArgvFormat.Number (the R8 locale rule).
The numeric tables are deliberately INTERNAL constants, not VideoCompileStepConfig knobs:
exposing raw eq/colorbalance numbers as workflow-author config would recreate, one layer up,
exactly the free-numeric-parameter surface the words-only contract closes off, for no capability
the curated looks don't already deliver. A future need for custom looks should add a NAMED look
to the table, not a numeric pass-through. (Considered and rejected: per-look config overrides
mirroring MusicBed*Db — those music knobs parameterize a level within one fixed mixing
topology, while grade numbers ARE the look itself.)
Compile integration
- One hard failure, up front:
EnableColorGrade=truewithMode=StreamCopyfailsCOLOR_GRADE_REQUIRES_REENCODE(a pure config error, mirroringGRAPHICS_REQUIRE_REENCODE). Everything else is soft:no_plan_configured,plan_unresolved,plan_invalid_json,unknown_look_word,look_none(the plan's own first-class no-grade decision, reported distinctly from every failure reason),grade_filters_unavailable(a probed, process-lifetime cached-filterscheck forcolorbalance/colorlevels/hue/eq, mirroring the drawtext/amix/xfade/perspective probes), and the defensiveempty_filter_chain— all degrade to "no grade applied", never a failed compile. - Stage ordering: grade before graphics. On the single-source select path the chain is
appended to the cut stage itself (directly after
setpts); on the segmented path it is its own stage directly after the concat/transition stage ([vcat]…[vgrd]). Either way it runs BEFORE any insert/overlay stage, so motion graphics always paint clean on top of graded footage — and unlike screen inserts, the grade works on BOTH encode paths (multi-source and transition-overlap compiles included), since it is a plain per-frame filter with no timeline bookkeeping to remap. - Reporting: the EDL and the step's output summary carry a
colorGradenode (enabled/applied/look/normalizedstrength/shadowTone/highlightTone/filterChain/reason) only whenEnableColorGrade=true—false(the default) leaves both shapes byte-identical to the pre-grade compile path, the same load-bearing guarantee asEnableGraphics/EnableMusic/EnableInserts.
AgentType.Colorist: the solo decider
An ordinary LLM Agent step (OutputSchemaName = "ColorGradePlanOutput"), minimal read-only tool
scope (ProjectRead + WorkflowControl — identical to VideoStoryEditor/MusicSupervisor;
nothing about a grade ever needs rendering, so no sandbox grant in any role). It reads the
bounded VideoAnalyze view's MEASURED per-shot words — the Phase 1/Phase 4 colour
temperature/tone/saturation/exposure descriptors and look groups — and picks the words above,
grounding its prose Reason in shot ids (s2, s4). Reasoning disabled, temperature 0.3, same
rationale as the other bounded pick-from-fixed-lists deciders.
The color grade room
StepType.ColorGradeRoom — the third room on
the shared room infrastructure: several Colorist-role seats
plus a supervising colorist (AgentType.ColorGradeDirector) deliberate in a live group chat over
the same bounded view a solo colorist consumes, then the director synthesizes ONE
ColorGradePlanOutput OUTSIDE the chat loop. What is specific to this room:
- The decision carries no ids at all — one whole-program grade in words.
FilterToOfferedIdsis a documented structural no-op (pinned by the invariant tests' property-set check), androom.droppedSpanCountis always0. The offereds{n}SHOT ids anchor only the DELIBERATION: seats argue per-shot ("s2 reads backlit", "s0 and s4 disagree on temperature"), and shot-id mentions drive convergence detection (ColorGradeRoomGroupChatManager, patterns\d+— deliberately narrower than the edit room's[sgt]so a gap/segment mention never counts, and never matchingp{n}/m{n}/k{n}). - Word validation replaces id filtering. Synthesis hard-rejects an EMPTY
Look(retry with the allowlist spelled out — "no grade" must be said as the wordNone, never as silence); a non-empty unknown look spends one retry viaCheckRetryableIssueand on the final attempt passes through to the compile step's ownunknown_look_worddegrade. Look: "None"is a fully valid outcome (the graphics room's empty-plan grace, in this room's vocabulary), and a view offering ZERO shots completes immediately with aNoneplan (room.terminationReason = "empty-view") instead of failingVIEW_UNRESOLVED.- Seats (
ColorGradeRoomStepConfig.DefaultSeats):ToneArtist(exposure/contrast, crushed or washed shots, faces),PaletteArtist(temperature/saturation words, which named look the footage wants),ContinuityArtist(one grade must suit every kept shot; whenSubtlebeatsStrongand whenNoneis right). All three resolve to the built-inColoristagent — personas injected per-turn, the established one-agent-many-personas pattern. - No tool-restriction override is needed — unlike the graphics room, both backing agent
types are minimal read-only by
ToolGroupCatalogin every role, so the base default (the agent's full grant) is already the restricted set.ColorGradeDirectortherefore mirrorsVideoEditDirector, notMotionGraphicsDirector. - Failure codes:
COLOR_GRADE_ROOM_CONFIG_INVALID/VIEW_UNRESOLVED/COLOR_GRADE_ROOM_FAILED/UNEXPECTED_ERROR. Solo fallback is one ordinaryAgentType.Coloristcall (FallbackToSoloColorist, defaulttrue).
ColorGradeRoomStepConfig (color_grade_room_config_json, jsonb on BOTH DbContexts) carries the
same knobs as EditRoomStepConfig/GraphicsRoomStepConfig (see
Config reference) with FallbackToSoloColorist as its fallback name. The
step's AgentDefinitionId FK is satisfied by the same AgentType.VideoTransform placeholder as
the other room/deterministic step types.
The video-derush-edit-grade-room template
An eighth opt-in template (AutoCreateOnProject: false): VideoAnalyze (Source: ProjectFile —
the default AnalyzeVisuals: true already emits the measured per-shot colour words, no extra
flag needed) → Agent(VideoStoryEditor) → ColorGradeRoom (View: Step 1 — explicit, since
Previous would resolve to the story editor's decision, not the measured-shot envelope) →
VideoCompile (Decision: Step 2, AnalysisStepOrder: 1, enableColorGrade: true,
colorGradePlan: Step 3) → ReviewLoop(VideoReviewAgent) looping back to step 2.
Deserialization-tested in WorkflowTemplateCatalogConfigDeserializationTests.cs.
Explicitly not built (color grading)
- Per-shot or per-look-group grades. The v1 grade is whole-program by design: per-shot
grading would need timeline-enabled filter windows remapped through every cut/transition (two
disagreeing sources of timing for one filter), and per-look-group grading would mean offering
the
k{n}namespace to an agent — a namespace this feature deliberately keeps descriptive-only. A future per-scope phase should anchor to offered ids, not timestamps, exactly as overlays did. - Numeric grade knobs in config — see the
ColorGradeFilterBuilderrationale above. - LUT files. A
.cubeupload path would reintroduce an opaque binary asset into the filter chain with no words-only audit trail; named first-party looks keep the EDL self-explanatory.
Sound effects
Optional discrete sound-effect cues — whooshes, clicks, dings, transition stingers, UI sounds —
mixed over the compiled edit's audio during VideoCompile's encode, each fired at the moment an
offered cut-anchor item begins in the OUTPUT timeline. The same structural shape
Background music established: a deterministic candidate list from
VideoAnalyze, an agent that chooses among opaque offered ids plus a handful of enum words, and
VideoCompileStepExecutor alone resolving those words to real ffmpeg behavior. Off by
default (VideoAnalyzeStepConfig.OfferSfxClips = false, VideoCompileStepConfig.EnableSfx = false) — EnableSfx = false leaves the compile path byte-identical to the pre-SFX behavior.
Why SFX is not music with a different name
Music and SFX share the "pick among uploaded audio/* files by opaque id" mechanics but are
deliberately DIFFERENT shapes, because the use cases are inverses of each other:
| Background music | Sound effects | |
|---|---|---|
| Cardinality | At most ONE track for the whole program | Zero or more discrete cues |
| Duration | Continuous bed, loops/fades to fit the edit | Short one-shot, hard-capped by MaxSfxCueSeconds (default 4s) |
| Placement | Program-wide — no moment to choose | THE decision: an offered cut-anchor id per cue |
| Ducking | Keyframed speech-envelope ducking under dialogue | None, deliberately — see below |
| No-agent path | MusicTrackProjectFileId (a fixed bed is a complete feature) | None, deliberately — see below |
No ducking for cues: a cue is a short, deliberately-audible accent — ducking a 300ms whoosh
under dialogue would defeat its purpose, and a duck/lift envelope per cue would multiply
filtergraph complexity for negative benefit. Loudness control is the Volume word plus
conservative default gains (SfxSubtleDb/SfxNormalDb/SfxStrongDb, defaults -18/-12/-6 dBFS,
clamped [-40, 0]), and the sound-designer prompt steers toward Subtle/Normal over dialogue.
No deterministic no-agent config path (rejected alternative): music's
MusicTrackProjectFileId works because "one fixed bed under everything" is a complete feature
with zero editorial judgment. The SFX equivalent would be "the same stinger at every cut" — a
well-known bad pattern that would ship as an attractive footgun, while any more selective
deterministic rule ("only section breaks", "only after long silences") smuggles editorial
judgment into config. Cue placement — WHICH clip at WHICH moment, and whether any moment
deserves one at all — is the whole decision, so the agent IS the feature here. A plan whose
cues list is EMPTY is a fully valid outcome (sfx.reason = "empty_plan"), the same first-class
no-op grace the graphics room's empty plan and the colorist's Look: "None" get.
Candidate discovery (VideoAnalyze)
When OfferSfxClips = true, VideoAnalyzeStepExecutor enumerates every audio/* project file
as an x{n} SFX-clip candidate (view.sfxClips, one {id, name} entry each), capped by
MaxSfxClips (default 40). Project-level, not per-source, and never ffprobed at analyze time —
the exact OfferMusicTracks rationale. When both OfferMusicTracks and OfferSfxClips are on,
ONE ListFilesAsync call feeds both candidate lists (and one listing failure degrades both to
zero candidates, never failing the step); the same file can legitimately appear under both an
m{n} and an x{n} id, since nothing structural distinguishes an uploaded stinger from an
uploaded bed — the sound-designer prompt tells the model to choose short one-shot clips by file
name and leave bed-like names to the music layer. view.sfxClips follows the exact
musicTracks budget discipline: shown as ONE ATOMIC ARRAY, not gated on VisualDetail, suppressed
as a whole (after musicTracks, before insertRegions) only once detail has fully degraded, and
always before any offered item is dropped. "Offered" means exactly "the whole array survived to
the final view" — VideoAnalysisArtifact.OfferedSfxIds is empty whenever it was suppressed.
x{n} id isolation — and the one namespace SFX deliberately shares
SFX-clip ids (x{n}) are their own namespace: never resolvable by
VideoCompileStepExecutor.BuildIdTimeIndex (an x{n} id names a FILE, never a moment), and a
Keep span naming one fails UNKNOWN_ID exactly like a placement/music/look id would. The
novel bit is the cue's ANCHOR: SfxCue.AnchorId deliberately IS a cut-anchor id
(s{n}/g{n}/t{n}) drawn from the same OfferedIds a Keep span may name — "this sound
fires when this shot/gap/segment begins" is the whole anchoring model, reusing the one id
vocabulary the model already reasons about instead of inventing per-moment SFX-placement ids.
Anchor validation is against OfferedIds (offered, not merely present in the artifact), and the
anchor's start is mapped through OutputTimeline.MapToOutputSec — an anchor whose moment was
cut away by the edit decision drops that cue (anchor_cut_away), never relocates it.
AgentType.SoundDesigner and SfxPlanOutput
An ordinary solo StepType.Agent step — deliberately NOT a room. The edit/graphics/grade rooms
exist because those decisions have genuine multi-perspective tension across a whole program (one
grade must suit every shot; overlays compete for placements and attention). An SFX cue is a
small, local decision over a small candidate list, made a handful of times; multiple
sound-designer personas deliberating each whoosh would multiply model calls for no perspective a
single well-prompted pass lacks. If a room ever proves warranted, the shared
RoomStepExecutorBase infrastructure makes it an additive follow-up, not a rewrite.
public class SfxCue
{
public string SfxId { get; set; } = ""; // must be in OfferedSfxIds; e.g. "x0"
public string AnchorId { get; set; } = ""; // must be in OfferedIds; e.g. "s2"/"g1"/"t3"
public string Timing { get; set; } = ""; // OnCut | Lead | Lag — never a number
public string Volume { get; set; } = ""; // Subtle | Normal | Strong — never a dB number
public string Reason { get; set; } = "";
}
public class SfxPlanOutput
{
public List<SfxCue> Cues { get; set; } = new(); // empty = a valid "no effects" decision
public string PlanRationale { get; set; } = "";
}
The rushcut invariant, extended a fifth time: every property is a plain string, pinned by
SfxPlanOutputInvariantTests (no numeric/time-bearing property anywhere; every SfxCue property
a string; both property sets pinned exactly, so a future OffsetMs-style addition fails CI).
Same minimal read-only tool scope as VideoStoryEditor/MusicSupervisor — every cue plays an
EXISTING uploaded clip, so there is no rendered-asset escape hatch and no sandbox grant in any
role. Reasoning disabled, temperature 0.3, same rationale as the other bounded
pick-from-offered-lists deciders.
Resolution: ResolveSfxAsync, soft-failure throughout
Runs only when EnableSfx = true. Two HARD config errors up front (SFX_REQUIRES_REENCODE,
SFX_REQUIRES_AUDIO_REENCODE — the exact music pair); everything else degrades:
| Situation | Outcome |
|---|---|
No SfxPlan configured / unresolvable / not valid JSON | No SFX (no_plan_configured / plan_unresolved / plan_invalid_json) |
Plan's cues list is empty | No SFX, reported as the VALID empty_plan outcome, distinct from every failure |
| Output has no dialogue and no applied music bed | The cue is mixed over a synthesized silent base — see Program audio on silent footage. (no_base_audio used to drop every cue here, and the effects were never heard.) |
This ffmpeg build's amix has no normalize option | sfx.unavailable = true, no SFX (the same cached probe music uses — normalize=0 protects the dialogue level) |
Cue's SfxId not in OfferedSfxIds / AnchorId not in OfferedIds | That cue dropped (unknown_sfx_id / unknown_anchor_id), the rest proceed |
| Anchor's start maps to no output moment (cut away) | That cue dropped (anchor_cut_away) — never relocated |
Clip not in project / not audio/* / download or probe fails | That cue dropped (sfx_not_in_project / sfx_not_audio / sfx_download_or_probe_failed) |
More cues than MaxSfxCues (default 8) | Excess dropped in plan order (max_cues_exceeded) |
Word resolution mirrors music's exactly: unknown Timing/Volume words normalize to
OnCut/Normal. Timing shifts the cue by SfxLeadMs/SfxLagMs (defaults 150ms) — applied
ON THE OUTPUT TIMELINE, after the anchor mapping, so a shifted cue can never land inside a cut
region its anchor's own moment survived — then clamps to [0, TotalSec]. Each cue's play window
is min(clip duration, MaxSfxCueSeconds, remaining output) — a long file misused as a cue is
trimmed, never allowed to become a de-facto bed — with a SfxFadeOutMs declick fade at its end
(capped at half the window). Distinct clips are downloaded/probed ONCE even when several cues
replay them (each cue still gets its own ffmpeg input, since each branch trims/gains/delays
independently).
The mix: SfxMixFilterBuilder
Pure static filter-fragment builder (Services/Video/SfxMixFilterBuilder.cs, exact-string-tested
by SfxMixFilterBuilderTests — the MusicMixFilterBuilder role). Each cue becomes one extra
-i input — always the LAST inputs, after every asset-overlay/screen-insert/music input, so
every existing input-index mapping (and every filter-string test asserting one) stays untouched —
and one branch:
[N:a]atrim=end={dur},asetpts=N/SR/TB,aformat=…48000…stereo,volume={gain},afade=t=out:…,adelay={ms}|{ms}[sfx{k}]
adelay (integer milliseconds, computed entirely server-side from the anchor mapping) is the one
and only place a cue's timing enters the filtergraph. The mix stage layers every cue over the
pre-SFX audio chain — which ends at an internal [abase] label when cues are present (the exact
[aout]→[adial] label-flip convention music established), whether that base is the plain
dialogue cut, the dialogue+music amix, or a music-only branch:
[abase][sfx0][sfx1]amix=inputs=3:duration=first:dropout_transition=0:normalize=0[aout]
Same three load-bearing amix options as music: normalize=0 (never quietly divide the
dialogue's level), duration=first (the base pins the output length — a cue near the end can
never extend the file), dropout_transition=0 (no gain re-ramp when a cue's short branch ends
early, which every cue's does). The SFX mix runs AFTER the music mix and BEFORE the seam-ramp/
program-fade tail stage, so an end-of-program cue fades out with the program envelope. Unlike
screen inserts (single-source-only in v1), cues work identically on BOTH encode paths —
multi-source and crossfade-overlap compiles included — since adelay against the output
timeline has no per-span bookkeeping to disagree with.
EDL / output JSON shape
Present only when EnableSfx = true (byte-identical to the pre-SFX compile path otherwise):
{
"sfx": {
"enabled": true, "applied": true, "unavailable": false,
"appliedCueCount": 2,
"cues": [
{ "sfxId": "x0", "anchorId": "s4", "clipName": "whoosh.wav", "timing": "OnCut",
"volume": "Normal", "gainDb": -12, "outputStartSec": 12.4, "playDurationSec": 1.8 }
],
"dropped": [ { "reason": "anchor_cut_away", "sfxId": "x1", "anchorId": "s2" } ]
}
}
The video-derush-edit-sfx template
A ninth opt-in template (AutoCreateOnProject: false): VideoAnalyze (Source: ProjectFile,
OfferSfxClips: true) → Agent(VideoStoryEditor) → Agent(SoundDesigner) → VideoCompile
(Decision: Step 2, AnalysisStepOrder: 1, EnableSfx: true, SfxPlan: Step 3) →
ReviewLoop(VideoReviewAgent) looping back to step 2. Same explicit-StepOrder rationale as
video-derush-edit-music (Previous relative to the compile step would resolve to the
SoundDesigner step's own output, not the story editor's decision). Deserialization-tested in
WorkflowTemplateCatalogConfigDeserializationTests.cs like every other template.
Explicitly not built (sound effects)
- A deterministic no-agent cue path — see the rejected alternative above.
SFX-only audio for silent outputs— now built: a silent base is synthesized, see Program audio on silent footage.- Per-cue ducking/sidechaining — see "no ducking for cues" above.
- Beat-matching, auto-selected libraries, generated/synthesized effects — every cue plays an uploaded project file the model was offered by id, nothing else.
Generated clips (b-roll)
A ShotDirector agent plans generated shots from a VideoAnalyze view; the deterministic
VideoGenerate step buys one clip per shot; VideoCompile places them in the edit. This is the
video-generation feature's side of the pipeline — docs/video-generation.md owns the generation
side (the agent's prompt, the provider layer, the spend ledger) and this section owns what the
compile side does with the result.
The template is video-derush-edit-broll (opt-in, AutoCreateOnProject: false):
VideoAnalyze → Agent(VideoStoryEditor) → Agent(ShotDirector) → VideoGenerate (plan-sourced)
→ VideoCompile (enableGeneratedClips) → ReviewLoop(VideoReviewAgent). Every step reference in
it is an explicit StepOrder, never Previous: Previous relative to the compile step would
resolve to the VideoGenerate step's own output rather than the story editor's decision — the
same rationale the music/SFX templates carry.
The planner's vocabulary is words and ids only
Each planned shot is purpose (Cutaway/SeamBridge/Extend/ColdOpen/EndCard/
ScreenContent), anchor (an offered s{n}/g{n}/t{n} cut-anchor id), optional
firstFrame/lastFrame (offered s{n} shot ids), camera (an enum word), duration
(Short/Medium/Long), prompt (free prose) and reason. There is no number, no URL, no
timestamp and no duration in seconds anywhere in it, pinned by a reflection invariant test. Every
numeric thing that reaches a provider — the snapped duration, the aspect ratio, the camera clause
appended to the prompt — is resolved by first-party C# in ShotPlanParameterMapper.
The prompt is the only field that leaves the building, so it goes through the shared
GeneratedPromptSanitizer (control characters stripped, per-provider length cap reported rather
than silently truncating) before it reaches a provider. A prompt that sanitizes to nothing, or
that exceeds the cap, degrades that clip with INVALID_PROMPT — the step continues.
There is no per-purpose placement config
Placement derives from each planned shot's own purpose, exactly as graphics and inserts derive
from their plan's own items. A second config surface would only let the two disagree. The compile
config gains exactly three fields, appended at the end of the positional record:
ExtractInputRef? GeneratedClips = null, // which step's resolved generated clips to place
bool EnableGeneratedClips = false, // false = byte-identical to the pre-wave compile path
int MaxGeneratedClips = 8);
EnableGeneratedClips: false — the default — is byte-identical to the pre-wave compile path:
no extra ffmpeg inputs, no extra filters, no new EDL or output JSON key. That is asserted on the
full argv sequence, not on the filtergraph string alone.
Clip id → storage key: the manifest, and why no path is model-visible
A generated step produces up to MaxClips clips, and every one of them must be addressable by the
compile side. WorkflowStepResult.OutputStorageKey still holds only the first succeeded clip (it
drives the player in the UI), so the per-clip manifest is the real contract:
projects/{projectId}/agentFiles/video-analysis/{executionId}/step-{order}-generated-clips.json
written by VideoGenerateStepExecutor and referenced from WorkflowStepResult.ArtifactStorageKey.
It carries clips[].{clipId, shotIndex, take, purpose, anchor, firstFrame, lastFrame, camera, duration, durationSeconds, status, storageKey, costUsd, costBasis, keyframeRequested, keyframeApplied, keyframeReason, failureReason, detail, skipReason} plus the step's meta, and
VideoCompileStepExecutor resolves config.GeneratedClips → that step's ArtifactStorageKey →
the manifest, the same way it already resolves the analysis artifact and the editorial decision.
Sitting under the video-analysis/ prefix rather than outputFiles/ means the existing
StepResultArtifactsController serves it with no controller change.
No storage path is model-visible. The step's OutputJson reaches later agents under
AgentInputContextMode.FullWorkflow, so it must not carry an object key: the manifest's
storageKey is projected out of Output and replaced by a stored boolean. That is the whole
reason the manifest exists as a separate artifact instead of an extra clips[].storageKey field,
and it is asserted directly — one test proves the Output contains no storageKey and none of
projects/, agentFiles/ or outputFiles/, while the manifest behind ArtifactStorageKey
carries the real keys.
Clip ids are v{n} (1-based plan order) when Takes is 1 and v{n}t{k} otherwise, so wave C's
dailies selection has N addressable take ids to choose between. The Inline prompt path keeps
phase 1's clip-{n} ids.
The placement table
purpose | Wave A behaviour |
|---|---|
Cutaway | Replaces picture only over the anchor's OUTPUT window — dialogue audio kept (J/L-cut) |
ColdOpen | Prepended as a concat segment, pairing with ProgramFadeIn |
EndCard | Appended as a concat segment, pairing with ProgramFadeOut |
SeamBridge, Extend, ScreenContent | Not applied in this wave (waves B/C) — degrade explicitly |
Cutaway resolves the anchor id through BuildIdTimeIndex + OutputTimeline to its window on
the output timeline, then applies the clip as an overlay on the program body's own video chain,
gated by enable='between(t,start,end)' with eof_action=pass and shortest=0. The window is
[anchorStart, min(anchorStart + clipDurationSeconds, anchorEnd)). Not one audio branch
references a generated input — no -itsoffset, no audio shift, no shortened stream — so the
anchor's dialogue plays continuously across the cutaway. An anchor whose start the edit cut away
degrades that clip with anchor_not_in_output; an empty window with empty_window.
A cutaway routes the compile onto the segmented encode path even for a single source. That is the same construct a single-source crossfade compile already produces; the alternative was two cutaway implementations that can drift.
ColdOpen/EndCard are normalized to the canonical canvas and added as real concat segments
(ColdOpen before segment 0, EndCard after the last), with matching-duration anullsrc
synthesized when the generated clip genuinely has no audio track — probed, never assumed. The
generated concat is the last content stage, immediately before the whole-piece program fade, so
ProgramEnvelopeFilterBuilder resolves against the final total including both extra segments
and the fades land on them. A cold open or end card that is present but unfaded because the
envelope was computed before it was added would be a bug.
Everything else degrades per clip, never failing the step — the compile is never taken hostage,
the same discipline graphics/music/inserts/grade/SFX already follow. A manifest that cannot be
resolved at all degrades the whole feature softly (generatedClips.reason = "plan_unresolved");
the one hard config error is EnableGeneratedClips with Mode=StreamCopy
(GENERATED_CLIPS_REQUIRE_REENCODE), because a picture replacement and a prepended segment both
need a filtergraph. At most MaxGeneratedClips clips are applied, in manifest order; the excess
degrades with max_generated_clips_exceeded. The result is reported in a generatedClips node on
both the output JSON and the EDL, present only when the feature is enabled.
Egress, and what this wave cannot yet do with a keyframe
A keyframe is not metadata about the user's footage — it is the footage. A shot naming
firstFrame/lastFrame therefore needs AllowSourceMediaEgress: true, evaluated per shot by
GeneratedShotEgressPolicy. A plan that asks for a frame without that consent fails as a whole
step with EGRESS_REFUSED and the policy's reason: zero clips are submitted and zero ledger rows
are written. A partially generated plan is a half-edit the operator cannot use, and the refusal has
to be actionable (grant consent or edit the plan) rather than silently dropping one shot.
With consent granted the step proceeds. This wave cannot attach the keyframe to the request:
VideoGenerationRequest has no image field, and image inputs arrive with phase 3's provider work.
So a consented keyframe-requiring shot is submitted best-effort and each such clip records
keyframeRequested: true, keyframeApplied: false,
keyframeReason: "keyframe_transport_not_available", counted in meta.keyframesNotApplied. It is
reported, never silently dropped. When a shot names a frame the config's Analysis reference is
required, and the named id is checked against the analysis artifact's offered shot ids — an id that
was never offered is recorded per clip as keyframe_unresolved, which is not fatal.
Budget: two estimates, one gate
MaxSpendUsd is checked against the whole plan's estimate before anything is reserved: the
pre-submit estimate sums Describe(model).PricePerSecond × snapped duration × takes across every
planned shot, and exceeding the cap yields BUDGET_EXCEEDED with zero submits and zero ledger
rows. That check is deliberately independent of the budget gate, so it holds even when MaxClips
trims the plan to a subset that would itself have been affordable — the cap is a ceiling on the
plan the author wrote, not on whatever survives the trim. The gate call that follows covers the
effective (post-trim) plan, and meta.planEstimatedUsd is that effective number, equal by
construction to the sum of the reservations handed to the gate. The project and global daily
budgets stay owned by VideoGenerationBudgetGate, unchanged.
Not built by wave A
SeamBridge,ExtendandScreenContentplacements (waves B and C) — the vocabulary exists, the placements degrade explicitly.- Keyframe transport to the provider (phase 3), and take selection over the
v{n}t{k}ids (phase 4). MaxClipsis still validated to 1..4, phase 1's bound, so a plan of more than four shots is trimmed rather than fully generated. Widening it is a cross-department config change (the builder mirrors the bound) and belongs with the dailies work.- A cold open or end card does not extend the background-music bed, the motion-graphics overlays or the tracked inserts over itself: those cover the program body. That is what keeps every overlay window and every SFX cue body-relative, so a prepended segment cannot silently shift them.
Builder UI coverage
Every StepType the backend knows is selectable and configurable in the workflow builder — the
step-type Select, the "Add Workflow Step" picker, the flowchart node map and the linear
(phone) step list all cover the same eleven types, and the picker's metadata is keyed as a
Record<StepType, …> so a future step type cannot compile without a UI entry.
The three room types share ONE editor and ONE canvas node
(web/components/workflows/RoomStepConfigEditor.tsx, nodes/RoomNode.tsx), mirroring the way
RoomStepExecutorBase<TDecision> backs all three server-side: ROOM_STEP_KINDS is the single
table holding each room's config field (editRoomConfigJson / graphicsRoomConfigJson /
colorGradeRoomConfigJson), its differently-named solo-fallback flag
(fallbackToSoloEditor / fallbackToSoloPlanner / fallbackToSoloColorist), and its copy and
accent. The editor exposes the fields that decide pipeline SHAPE (which analysis view the room
deliberates over, rounds, turn ceiling, the two degrade toggles) and spread-preserves every other
field, so a template-provisioned config (seats, temperatures, termination mode, timeouts) still
round-trips losslessly through a save. A freshly-added or type-switched room step is written with
a runnable default config immediately, since the executor hard-fails on an empty one.
VideoAnalyze's panel covers the candidate-offering switches each downstream agent needs
(EmitOverlayPlacements, OfferMusicTracks, OfferSfxClips, DetectInsertRegions with its
allowlisted plate colour) plus vision captioning; VideoCompile's covers every optional
post-production stage (EnableGraphics, EnableInserts, EnableColorGrade, EnableMusic
including the deterministic MusicTrackProjectFileId path, EnableSfx) and their plan
references. Each of those toggles switches Mode back to Reencode when turned on from
stream-copy, since every one of them is documented as requiring a re-encode — the backend would
otherwise reject the step as a config error at execution time.
Deliberately still raw-JSON-only: per-seat persona lists, sampling temperatures, reasoning effort, room timeouts, and the numeric trim on graphics/music/SFX (per-word dB tables, overlay box geometry, lift-window shaping). Those are tuning, not pipeline shape.
Security: why ffmpeg is not in the sandbox
ffmpeg and ffprobe run as first-party, non-AI-authored C# code inside the WorkflowEngine
container. The Remotion sandbox service (/sandbox) — its allowlist, its images, its threat
model — is untouched by this feature.
The sandbox's command allowlist is narrow because the code running inside it is untrusted: it
executes model-authored TSX and installs model-chosen npm packages. Admitting ffmpeg to that
allowlist would open one of the richest argv-injection surfaces in common Unix tooling —
-i http://… (SSRF into the sandbox's bridge network, which is a normal routable Docker network,
not --network none — see docs/sandbox-service.md), the concat:/subfile:/file: protocols
(arbitrary in-container file read), -f lavfi with movie=, arbitrary output paths, arbitrary
-map. Making that safe requires an argv validator at least as strict as simply constructing the
argv yourself server-side — at which point the sandbox's containment has bought nothing, and a
hardened boundary has been widened for free.
Meanwhile, the containment property the sandbox exists to provide is not needed here: this feature's ffmpeg argv is built entirely by first-party C# from a validated, typed cut list. The model's only contribution is a set of opaque ids drawn from a set the system itself issued — not one model-originated character reaches an ffmpeg argv.
Qualification (Phase 3): motion-graphics overlay TEXT is model-authored and does reach ffmpeg — but never as argv or filter-string content. It is sanitized to an allowlist, written to its own scratch file, and referenced only via drawtext's
textfile=option (withexpansion=noneset as defense-in-depth). See Motion graphics (Phase 3) for the full discipline. The claim above — "not one model-originated character reaches an ffmpeg argv" — still holds exactly as stated for the argv/filter-string surface; it is the reason Phase 3's text still cannot inject into it.
Same boundary, the render side (B8b): the trial output policy's 720p cap and ReelBolt watermark are applied to a Remotion render by a POST-PASS in the engine — the render is pulled out of the sandbox, re-encoded once by
ReactRemotionSandboxToolsthrough the sameOutputPolicyFilterBuilderchain the compile step splices in, and uploaded from the engine. The filter string is first-party C# built from two integers and a configured asset path; the model's render output never contributes an argv character here either. The sandbox still has no ffmpeg.
The sandbox's mechanics are also concretely wrong for large binary media: containers run
--read-only with a 256 MB tmpfs /tmp; sandbox file I/O is base64-over-JSON (a 400 MB mp4
becomes a ~533 MB base64 string materialized in .NET memory in both directions); sandbox
containers cannot reach the object store, so a video routed through the sandbox would round-trip through the
engine anyway; and sandbox container lifetime is keyed to workflowExecutionId and janitor-TTL'd
— the wrong lifecycle for a step that must produce a durable artifact.
Costs of running ffmpeg in-process, and how they're mitigated:
| Cost | Mitigation |
|---|---|
CPU-heavy encoding competes with WorkflowEngine:MaxConcurrency | A singleton SemaphoreSlim in IVideoToolRunner; VideoEditing:MaxConcurrentJobs defaults to 1. The semaphore wait is cancellable and excluded from the ffmpeg timeout. |
| ffmpeg parses untrusted, user-uploaded media (a native attack surface) | -nostdin -hide_banner -y -protocol_whitelist file on every invocation (-loglevel error for encoding; raised to info only for silence/shot detection, since those tools log their markers at ffmpeg's info level and the detector parses them from stderr); every input path is asserted to be under the per-execution scratch dir; pre-decode caps on byte size and probed duration; the process runs as the container's existing non-root $APP_UID. |
| Zombie processes on cancel/shutdown | ct.Register(() => proc.Kill(entireProcessTree: true)) plus a hard wall-clock timeout. |
| Container image grows (ffmpeg + its dependencies) | Accepted as the cost of this design; ffmpeg is an Alpine package, not a large custom build. Measured (WS7): the built workflow-engine image is ~183 MB larger than the otherwise-identical inference image built from the same base (mcr.microsoft.com/dotnet/aspnet:9.0-alpine) — mostly ffmpeg's own codec dependency tree (libx264, libx265, libvpx, libaom, libsvtav1, vulkan loader, etc.), not the ffmpeg binary itself. |
The documented phase-2 escape hatch, if this ever needs to scale independently or run with
different trust boundaries, is a dedicated video-worker microservice mirroring the sandbox's
per-job container model — correct long-term shape, disproportionate for a first iteration. All
ffmpeg invocation sits behind a single IVideoToolRunner interface specifically so that swap is a
one-class change later.
Explicitly not built
Editing scope: no reordering of kept spans (v1 requires strictly increasing, non-overlapping
spans within one clip); no arbitrary unanalyzed asset/image insertion (cutting across several
pre-declared, analyzed Sources clips — including B-roll — is supported, see
Multiple source clips, but inserting an image or a clip that was never
fed in as a Source is not); no multicam (no automatic multi-angle sync/switching); no
free-floating picture-in-picture (tracked screen inserts ARE built — a rendered scene composited
into a tracked chroma-plate region, see
Tracked screen inserts (Phase 5) — but an arbitrary
un-tracked inset window is not); no speed ramps; no agent-requested or agent-authored transitions — only the
deterministic, measurement-driven seam treatments and program fade described in
Seam transitions and the program envelope.
Post scope: no colour grading / LUTs / filters / stabilization (Phase 4's D1-D3 color
dimensions are measured/reported only, same as loudnorm below, never applied); no loudness
normalization (loudnorm is measured and reported only, never applied); no subtitle burn-in and no
SRT/VTT export (the transcript exists, so this is the most obvious phase-2 add); no speaker
diarization. Background music IS built — see Background music — but it is a
deterministic bed/ducking mix only: no auto-composed score, no beat-matching to cuts, no per-section
music cues. Discrete sound-effect cues ARE built — see Sound effects — but only
as agent-planned, id-anchored playback of uploaded clips: no generated/synthesized effects, no
deterministic cue-at-every-cut path, no per-cue ducking.
Interchange: no EDL/AAF/FCPXML/OTIO export. The internal EDL JSON is an audit artifact, not an interchange format.
Timecode: no drop-frame handling, no SMPTE timecode parsing or emission, no timecode tracks.
Remotion's own template renders integer 30 fps h264/mp4, and an uploaded video's rational fps
(e.g. 30000/1001) is handled exactly via integer frame arithmetic without needing a timecode
subsystem.
Delivery: no streaming/HLS packaging, no proxy/preview transcodes, no thumbnail sprites, no waveform PNGs.
Approval: no suspend/resume human-approval gate mid-execution — "approval by composition" (run analysis + decision, inspect, then run compile separately against the saved artifact) instead. A real suspend/resume gate is the single highest-value phase-2 item.
Platform: no changes to /sandbox whatsoever — not the Go source, not the allowlist, not the
images, not the compose entry. No new microservice in v1. No async/polling execution model beyond
what the workflow engine already has. No ProjectFile backfill for historical render outputs.
Limits: MaxDurationSeconds defaults to 1800 (30 min), MaxInputBytes to 2 GB — revisit if
real inputs are much longer or shorter. No GPU encoder path (e.g. h264_nvenc) unless the
deployment host is confirmed to have one.