Skip to main content

Video Editing

Automatic derushing and editing of real, uploaded or rendered video files — silence and shot detection, optional ASR transcription, an LLM editorial decision, and a frame-accurate ffmpeg cut, extended with multi-clip cutting, motion graphics, background music, and a multi-agent "edit room" deliberation — implemented as three new deterministic-or-orchestrating workflow step types plus five new built-in LLM agents (VideoStoryEditor, MotionGraphicsPlanner, VideoReviewAgent, MusicSupervisor, VideoEditDirector) and one deterministic placeholder agent (VideoTransform). This document is the reference for that feature; for the surrounding workflow engine (step types, executors, agents in general) see CLAUDE.md.


Table of Contents​


The three-stage shape​

┌─────────────────┐ ┌────────────────────────┐ ┌──────────────────┐
│ StepType. │ │ StepType.Agent │ │ StepType. │
│ VideoAnalyze │────▶│ AgentType. │────▶│ VideoCompile │
│ │ │ VideoStoryEditor │ │ │
│ deterministic, │ │ LLM, structured output │ │ deterministic, │
│ ffmpeg + ASR │ │ (VideoEditDecision- │ │ ffmpeg │
│ │ │ Output) │ │ │
└─────────────────┘ └────────────────────────┘ └──────────────────┘
emits {view, meta} emits an ordered list of resolves ids to
+ a full analysis opaque "Keep" id spans frame-accurate
artifact — never a timestamp times, encodes
  1. StepType.VideoAnalyze (ReelBolt.Shared/Workflows/VideoAnalyzeStepConfig.cs, WorkflowEngine/Execution/StepExecutors/VideoAnalyzeStepExecutor.cs) — deterministic, non-LLM. Resolves the source video, probes it, detects silence gaps and shot/scene changes with ffmpeg, optionally transcribes the audio, assigns every detected item a short opaque id, and emits:

    • a bounded {view, meta} prompt envelope (the same shape StepType.Extract produces, so downstream Extract steps can compose over it unchanged), persisted to output_json
    • the full, non-truncated analysis document (VideoAnalysisArtifact), persisted separately as a storage-key artifact
  2. StepType.Agent + AgentType.VideoStoryEditor — an ordinary Agent step. No new step type was introduced for the editorial decision: AgentStepExecutor already provides structured output (ChatResponseFormat.ForJsonSchema<T>()), per-agent provider resolution, retry-with- feedback, and tool scoping, so reusing it is a straight win. The agent is given the bounded view from step 1 and decides which ids to keep.

  3. StepType.VideoCompile (ReelBolt.Shared/Workflows/VideoCompileStepConfig.cs, WorkflowEngine/Execution/StepExecutors/VideoCompileStepExecutor.cs) — deterministic, non-LLM. Loads the full analysis artifact from step 1, resolves the agent's chosen ids to exact [start, end) times, validates and normalizes the resulting cut list, frame-quantizes it, and encodes the edited video with ffmpeg. Optionally (Phase 3, EnableGraphics), also composites motion-graphics overlays during that same encode — see Motion graphics (Phase 3).

An optional fourth stage — StepType.Agent + AgentType.MotionGraphicsPlanner — can sit between steps 2 and 3, planning overlays from placement candidates step 1 derived alongside its usual cut anchors. It follows the exact same shape as step 2 (an ordinary Agent step, no new step type). See Motion graphics (Phase 3) for the full design.

Both VideoAnalyze and VideoCompile follow the same discipline ExtractStepExecutor established: they never throw. Every failure mode — an unresolvable source, a missing transcription provider, an unknown id, overlapping spans — is represented as a structured failure in the step's JSON output, because WorkflowStepResult.OutputJson is a jsonb column and an unhandled exception there takes down the whole execution's SaveChangesAsync.

Picking the source video​

VideoAnalyzeStepConfig.Source is a VideoSourceRef, one of three kinds:

KindResolves viaUse case
PreviousStepOutput (default)The latest completed step result in this execution with a non-null output storage key"Edit whatever the previous step in this workflow rendered" — e.g. chain straight off a Remotion render step
StepOutput + StepOrderThat specific step's WorkflowStepResult.OutputStorageKeyReference an earlier step explicitly, regardless of what runs in between
ProjectFile + ProjectFileIdproject_files.storage_keyEdit a raw video/audio file the user uploaded

This three-way split exists because Remotion render outputs are not ProjectFile rows — ReactRemotionSandboxTools.RenderVideoAndUploadToStorage uploads directly to projects/{projectId}/outputFiles/{executionId}/{name} and records only WorkflowStepResult.OutputStorageKey, never inserting into project_files. A source model that could only address ProjectFile ids would be unable to express "edit the video this workflow just rendered" — the primary use case.


The id-anchored decision contract​

The rushcut invariant: the story-editor agent can never emit a timestamp. This is not a prompt instruction the model could drift away from — it is structural. VideoEditDecisionOutput (ReelBolt.Shared/Data/OutputSchemas.cs) has no numeric, TimeSpan, or DateTime property anywhere in it:

public class VideoEditKeepSpan
{
public string FromId { get; set; } = ""; // e.g. "t7" or "s2" — first offered id to keep
public string ToId { get; set; } = ""; // last offered id to keep, inclusive
public string Reason { get; set; } = ""; // prose only
}

public class VideoEditDecisionOutput
{
public List<VideoEditKeepSpan> Keep { get; set; } = new(); // ordered, non-overlapping
public string EditRationale { get; set; } = "";
public string SuggestedTitle { get; set; } = "";
}

There is no "remove" list — everything not covered by a Keep span is cut. A second, "remove ids" representation would create two ways to express the same edit and a resolution order to get wrong, for zero benefit.

This is enforced in CI, not just by convention: VideoEditDecisionOutputInvariantTests (inference/tests/ReelBolt.WorkflowEngine.Tests/VideoEditDecisionOutputInvariantTests.cs) uses reflection to assert that neither VideoEditDecisionOutput nor VideoEditKeepSpan exposes any int/long/float/double/decimal/TimeSpan/DateTime/DateTimeOffset property (nullable variants included). If a future change adds e.g. a StartSec field "to make things easier", this test fails the build.

Ids, and why "offered" is a stricter check than "exists"​

Every shot, silence gap, transcript segment, and word in the full analysis artifact gets a deterministic id, assigned by index: s{n} shots, g{n} silence gaps, t{n} transcript segments, w{n} words (words are persisted for audit/phase-2 subtitle export but never offered to the model — segments are the offered granularity in v1). The artifact also records OfferedIds: exactly which ids were actually included in the bounded view shown to the model, which can be a strict subset of every id in the full artifact once VideoAnalyzeStepConfig's view-budget trimming drops trailing items.

VideoCompileStepExecutor rejects any FromId/ToId that is not in OfferedIds — not merely present somewhere in the full artifact. This closes a subtle hole: without it, a model could reference an id it was never actually shown (e.g. one dropped by trimming), which would resolve to a real time in the artifact but not one the model ever saw or reasoned about.

How compile resolves ids to frame-accurate times​

VideoCompileStepExecutor, in order — never trusting a model-emitted number, because there is none in the schema to trust:

  1. Load the full analysis artifact via the referenced VideoAnalyze step's ArtifactStorageKey (AnalysisStepOrder, or AnalysisStepResultId to resolve a prior execution's artifact — scoped to the same project only, see below).
  2. Reject any id not in OfferedIds (UNKNOWN_ID).
  3. Map each id to [StartSec, EndSec) from the full artifact.
  4. Normalize: sort by start; require strictly increasing, non-overlapping spans (reordering kept spans is out of scope for v1 — a violation fails the step with a precise diagnostic that becomes retry feedback); coalesce adjacent spans; apply PrePaddingMs/PostPaddingMs; clamp to [0, durationSec]; drop spans shorter than MinSegmentMs; enforce MaxSegments.
  5. Frame-quantize on the exact rational fps read from ffprobe's r_frame_rate (e.g. 30000/1001), using integer arithmetic throughout — never a rounded double — so accuracy never depends on the source's frame rate being a whole number.
  6. Evaluate Expect (including MinRetainedRatio, which refuses an edit that discards nearly everything the source contained).
  7. Write the EDL artifact (the resolved cut list, for audit), then encode.

AnalysisStepResultId lets a VideoCompile step reference a VideoAnalyze result from a different execution — the first-class mechanism behind the approval-by-composition workflow below. It is resolved only against step results whose WorkflowExecution.ProjectId matches the current execution's project; a cross-project reference is refused.

The cut's rationale is joined to its OWN Keep span, never by loop index​

TimelineItem.AgentRationale on each cut-N item is the Reason of the VideoEditKeepSpan that cut actually descends from. That join used to be by loop index i — decision.Keep[i].Reason, else the first non-empty reason anywhere in Keep — which is only valid while the raw Keep list survives the span pipeline 1:1, and step 4 above makes sure it does not: spans are coalesced, spans below MinSegmentMs are dropped, and the list is capped at MaxSegments. After any of those, index i names a different Keep entry than the one the cut came from, so a cut routinely showed another cut's reason, and a cut with no reason of its own silently inherited the first reason in the list.

Instead, every span now carries ResolvedSpan.OriginKeepFromIds: the FromId(s) of each Keep[i] it descends from, threaded through the raw-span tuple pipeline from the moment the span is first resolved (new[] { span.FromId }) to the frame-quantized ResolvedSpan — including the no-valid-fps early return, so a span keeps its origins even when quantization cannot run. CoalesceAdjacent concatenates both sides' origin ids in order (previous side first) rather than keeping one, so a coalesced span still names every Keep entry it covers. BuildVideoTrack builds a keepByFromId dictionary once, outside the loop, and resolves the rationale by looking each origin id up in it.

A span whose origin list is null or empty — and a span whose origin ids all resolve to a blank reason — falls back to decision.EditRationale, the edit-level rationale. The old index lookup and its "first non-empty reason anywhere" fallback are gone: there is no path left that can attribute one cut's reason to another. See The timeline manifest for the item fields this rationale lands in.

Approval: composition, not a suspend/resume gate​

There is no blocking mid-execution approval gate in v1 — WorkflowExecutorService.ExecuteAsync has no AwaitingApproval status or persisted suspend point to resume from; building one is a whole initiative of its own. Instead:

  1. Run a workflow containing only VideoAnalyze + the VideoStoryEditor agent step. Inspect the proposed edit in the execution UI (the "Edit decision list" panel, or the VideoAnalyze step's own bounded view/artifact).
  2. Run a second workflow or execution containing only VideoCompile, pointed at the first run's analysis artifact via AnalysisStepResultId.

The compiled video needs no new playback UI: writing OutputStorageKey under the outputFiles prefix makes it appear in GET /outputs and in the existing execution-page <video> player automatically, exactly like a Remotion render.


Where artifacts live​

ArtifactStorage keyReferenced by
Full analysis JSON (shots + silences + words + segments + OfferedIds + provenance — can be several MB)projects/{projectId}/agentFiles/video-analysis/{executionId}/step-{order}-analysis.jsonWorkflowStepResult.ArtifactStorageKey on the VideoAnalyze result
Bounded prompt view {view, meta} (≤ MaxOutputChars, default 24,000)WorkflowStepResult.OutputJson (jsonb)Agent input / step history, same as any other step
Resolved EDL (the compile step's own cut list, for audit)projects/{projectId}/agentFiles/video-analysis/{executionId}/step-{order}-edl.jsonWorkflowStepResult.ArtifactStorageKey on the VideoCompile result
Timeline manifest (the finished edit as named studio tracks + per-item AI provenance)projects/{projectId}/agentFiles/video-analysis/{executionId}/step-{order}-timeline.jsonWorkflowStepResult.TimelineStorageKey on the VideoCompile result (see The timeline manifest)
Edited mp4projects/{projectId}/outputFiles/{executionId}/{OutputFileName}WorkflowStepResult.OutputStorageKey and a new project_files row (category outputFiles, mime video/mp4)
Working files (intermediate segment files, caption keyframes){VideoEditing:ScratchPath}/{executionId}/{stepId}/ (default /var/tmp/reelbolt-video)Nothing — deleted in a finally block once the step completes
Extraction result, for resumeprojects/{projectId}/agentFiles/video-analysis/{executionId}/step-{order}-extract.v1.jsonNothing — the key is derived from (project, execution, step order) by VideoAnalysisExtractArtifact.BuildStorageKey
The extracted full-track WAVprojects/{projectId}/agentFiles/video-analysis/{executionId}/step-{order}-src{n}.wavVideoAnalysisExtractSource.AudioStorageKey, carried inside the extract document

ArtifactStorageKey is a separate column from OutputStorageKey on purpose. OutputsController.ListOutputs treats any step result whose OutputStorageKey starts with projects/{id}/outputFiles as a playable video, and the execution UI renders a <video src=…> for it. Putting the analysis JSON or EDL there would make the UI try to play JSON as a video. Keeping them in a separate column under a separate agentFiles/video-analysis prefix avoids that entirely, and lets a dedicated endpoint validate against a narrower, non-playable prefix:

GET /api/v1/projects/{projectId}/step-results/{stepResultId}/artifact

Mirrors OutputsController.DownloadOutput's guards exactly (project exists → caller owns it → step result resolved through WorkflowExecution.ProjectId) plus one more: the storage key must start with projects/{projectId}/agentFiles/video-analysis — never outputFiles — so this endpoint can never be coerced into serving a playable render or another project's object.

Registering the edited video as a ProjectFile (unlike Remotion render outputs, which are not) is deliberate: it makes an edited video re-editable via VideoSourceKind.ProjectFile and visible in the project's file list. It's created with SummaryStatus = Done and IndexingStatus = NotIndexed so the text summarizer/vector chunker never touches a binary — see the R22 fix below.

The EDL is written and uploaded exactly once per compile. VideoCompileStepExecutor builds edl.json, writes it to scratch and uploads it under its final key before it starts encoding; the uploaded document deliberately does not carry timelineStorageKey inline. An earlier version re-wrote that same local file and re-uploaded the same storage key after the manifest existed, purely to stamp that one field onto it — doubling the EDL's upload cost on every compile, for a value every consumer already had elsewhere, and making the "one EDL artifact per run" invariant untrue. The key reaches callers through the step's OutputJson (timelineStorageKey) and WorkflowStepResult.TimelineStorageKey only; nothing reads it from inside the EDL. The single write+upload is unconditional and runs before the encode, so the ENCODE_FAILED path still returns its edlStorageKey — a failed encode keeps the cut list describing what it tried to cut.

The persisted extraction (resume)​

A VideoAnalyze step runs in two phases: a deterministic extraction phase — stage, probe, silence, shots, the full-track WAV, audio levels, frame grids and everything read off them, chroma tracking, look grouping, candidate listing, beat facts, the caption keyframes — and an inference phase (ASR, vision captioning, subject location, object tracking, the derush pre-pass, input screening, the artifact and the bounded {view, meta} envelope) that is the only thing which calls a provider.

The extraction phase's result is persisted as a versioned document (analysis-extract.v1.json, VideoAnalysisExtractArtifact) at the key in the table above, and a later attempt of the same step loads it before doing anything else — one HEAD, then a GET. A hit skips extraction entirely, so an ASR outage, a vision provider timing out on the last of ten captions, or any other inference-side failure costs the inference phase and nothing else. That is what a remote target needs in order not to re-dispatch a runner's extraction because the LLM half failed.

What a resume recovers, and what it deliberately does not. The analysed media is re-staged from the object store key the document records (for a still image the rendered clip, for an author-trimmed source the trimmed window — the media every id and time in the document is relative to), and the extracted WAV from its published key. Both are downloads, never decodes. The caption keyframes are not published (publishing them would duplicate the user-visible PersistKeyframes output and collide with the -keyframes/ path that output owns), so a resumed attempt whose session is gone re-solves the keyframes it cannot find — one seek plus one frame each, bounded by MaxCaptionedShots — and records any that still fail exactly where extraction would have. A session that survived the failed attempt (a remote/runner session) finds them present and pays nothing.

The WAV publish is load-bearing for the resume. It stays non-fatal, because in place inference reads the session file and a transient upload error must not fail a step that would otherwise succeed — but a missing key now has a defined consequence instead of a silent one: a document whose source has an extracted WAV but no published key is refused, and extraction runs again. That is the honest answer, because the alternative is a resume that reads no audio and reports an empty transcript for a video that has speech.

The document never becomes a source of truth. It is addressed by (project, execution, step order), so "it is still there" is not the same question as "it is still valid": a loader re-resolves every source reference and refuses the document when an upstream output has changed (a ReviewLoop loop-back), when the schema version is not the one this build writes, or when the source count no longer matches. Every refusal degrades to extracting again, which is exactly the behaviour the step had before the document existed. Losing the document costs a re-extraction and never correctness, so a failure to persist it is logged and never fails the step.

The timeline manifest​

Every successful VideoCompile also emits a timeline manifest: a TimelineManifest (ReelBolt.Shared/Workflows/TimelineManifest.cs) describing the finished edit as named studio tracks (video, inserts, graphics, dialogue, sfx, music), each holding TimelineItems positioned on the output timeline. Its key is published on WorkflowStepResult.TimelineStorageKey — a third column, separate from OutputStorageKey and ArtifactStorageKey for exactly the reason above (the manifest is JSON, and the execution UI plays anything under outputFiles). It is indexed on workflow_step_results, so TimelinesController.ListTimelines is a single indexed predicate rather than a scan over OutputJson, and the value is mirrored onto workflow_step_cache_entries and restored onto the reconstructed result on a cache hit — a cached VideoCompile still lists its timeline.

This document covers the v1 manifest — a read-only description of a finished compile — and the pipeline that produces it; the v2 editable document (EditTimeline) a client opens, edits and saves back, together with its op catalogue, sync protocol and trust boundary, is docs/editor.md, and the manifest is what a v2 timeline is seeded from.

TimelineItem.AgentName/ProviderName/StepResultId are reported facts, never inferred. They name the agent that really authored the item's decision, the inference provider that agent's step resolved to, and that step's newest WorkflowStepResult id — each read from the execution's own step graph and database rows. TimelineProvenance/TimelineProvenanceContext (WorkflowEngine/Services/Video/TimelineProvenance.cs) splits the work: the executor resolves the facts (it owns the DbContext scope and the step graph), and TimelineManifestBuilder — which stays pure, synchronous and I/O-free — stamps them. The builder takes a single trailing, defaulted provenance parameter, so every pre-existing call site is unchanged.

Each track stamps its items from its own slot of that context, keyed off the same VideoCompileStepConfig reference the decision itself was resolved from:

Track (TimelineTrack.Type)Provenance slotConfig reference
video (cuts)DecisionDecision
graphicsGraphicsGraphicsPlan
musicMusicMusicPlan
sfxSfxSfxPlan
video (cuts) — fallback onlyColorGradeColorGradePlan
inserts, dialogue(no slot — never stamped)deterministic: no agent authored them

Cuts win over the colour grade. A video item takes the Decision provenance whenever it resolves; the ColorGrade slot fills in only when the decision provenance is unavailable and the item really carries a ColorGrade value. A colour grader did not author an ungraded cut, so stamping grade provenance onto one would be precisely the fabricated label this provenance exists to remove. The two are never merged: one item names exactly one author.

A room step resolves its director, not its placeholder. EditRoom/GraphicsRoom/ ColorGradeRoom steps point their own AgentDefinitionId at the non-LLM VideoTransform placeholder that merely satisfies the non-nullable FK, so stamping it would have labelled every room-authored cut VideoTransform. The real author comes from the room's own config surface, exactly as RoomStepExecutorBase resolves it: the configured DirectorAgentDefinitionId when present, else the room kind's built-in director type (VideoEditDirector, MotionGraphicsDirector, ColorGradeDirector). A room config that is present but undeserializable returns neither — whether an override was configured is then unknowable, so the track degrades to null rather than guessing the built-in director.

All three fields null is a real answer, not a defect. Provenance resolution is best-effort by contract and can never fail a compile: an unresolvable step reference, a step absent from the execution, a missing AgentDefinition row, or a service scope with no IInferenceProviderResolver all leave that track's items unstamped — exactly what the manifest looked like before provenance existed. The three facts are independent, so a partial answer is normal: a step result not yet persisted loses only StepResultId, and a provider-resolution failure loses only ProviderName, keeping the agent name. Null therefore means deterministic (inserts, dialogue) or unresolved (everything else), never "guessing failed so here is a plausible label". TimelineInspector (web/components/editor/TimelineInspector.tsx) reads it the same way — provenance is reported by the backend, never re-derived on the client from which other fields happen to be populated, so an item with no agentName simply shows no author rather than an invented one.

Uploaded video/audio never reaches the text pipeline (R22)​

ProjectFilesController.Upload (and the folder-move and bulk-reindex endpoints) skip both summarization-queue enqueue and vector-indexing enqueue when the file's MIME type starts with video/ or audio/, setting SummaryStatus = Done and IndexingStatus = NotIndexed immediately instead of the normal Pending → summarizer/chunker pipeline. Without this, a large uploaded mp4 or wav would be handed to ProjectFileIndexingConsumer, which reads the object as UTF-8 text before chunking it — wasted compute at best, a very large in-memory string at worst.


Config reference​

VideoAnalyzeStepConfig​

FieldDefaultNotes
Version—Config schema version
Source—VideoSourceRef — see Picking the source video. Ignored when Sources is non-empty
SourcesnullMulti-source addition — IReadOnlyList<VideoSourceRef>. When non-empty, the AUTHORITATIVE list of source clips analyzed into ONE merged artifact; null/empty (default) falls back to treating [Source] as a one-element list — see Multiple source clips
DetectSilencetrueffmpeg silencedetect
SilenceThresholdDb-34.0
MinSilenceMs350
DetectShotstrueffmpeg scene-change detection
SceneThreshold0.30
MaxShotSec0Split every detected shot longer than this into sub-shots carrying part — see Keeping part of a long shot. 0 splits nothing (byte-identical)
TranscriptionOptionalOff / Optional / Required — see Transcription
TranscriptionProviderIdnullExplicit override; otherwise resolved via the default Transcription-capability provider
LanguagenullPassed through to the ASR provider
WordTimestampstrue
MaxAsrChunkBytes20,000,000ASR chunk size cap (~10.4 min of 16kHz mono s16 WAV); long audio is split at silence-boundary-aligned chunks, never mid-word
MaxDurationSeconds1800Guardrail, checked before any decode
MaxInputBytes2,000,000,000Guardrail, checked before any decode
MaxOutputChars24,000Prompt-view budget — trimming drops whole trailing items and re-serializes, never truncates mid-JSON (same rule as ExtractStepConfig)
MaxViewSegments400
MaxSegmentTextChars160
AnalyzeVisualstruePhase 1: one low-res grid ffmpeg pass + pure C# analyzer — see Scene/visual analysis (Phase 1)
VisualSampleFps2.0Grid sample rate, clamped downward by MaxVisualSampleFrames for long videos
VisualGridWidth / VisualGridHeight32 / 18Downscaled grid resolution the analyzer runs against
MaxVisualSampleFrames4000Caps the grid buffer size: effectiveFps = min(VisualSampleFps, MaxVisualSampleFrames / durationSec)
StillMotionThreshold0.02Per-frame motion (0..1) below which a moment counts as "still"
MinStillWindowMs400Minimum duration for a still run to be reported as a StillWindow
MaxStillWindowsPerShot3Longest still windows kept per shot
DetectLetterboxtrueD6: implemented (Phase 4), free from the grid data Phase 1 already samples — see Semantic visual dimensions (Phase 4). Default changed from false; may under-report soft/gradient letterbox edges
DetectSharpnessfalseD-adjacent "sharpness": implemented (Phase 4) but costs one extra native-resolution ffmpeg invocation per measured shot (capped by MaxSharpnessShots), so it stays opt-in unlike the other free dimensions — see Semantic visual dimensions (Phase 4)
AnalyzeAudioLevelstruePhase 1: WavRmsSampler over the WAV already extracted for transcription, or extracted fresh if transcription is off
DetectNearDuplicatestruePhase 1: near-duplicate/best-take grouping via FrameGridAnalyzer.GroupDuplicates
DuplicateSimilarityThreshold0.90Minimum signature similarity (0..1) for two shots to be grouped
DuplicateWindowShots20Single-linkage grouping only compares a shot against the previous N shots (multi-take shots are temporally adjacent)
VisualDetailCompactNone / Compact / Full — how much per-shot visual/audio detail the bounded view includes; degrades toward None before any item is ever dropped — see below
MaxViewDuplicateGroups20Caps view.duplicateGroups
AnalyzeColorGradingtruePhase 4, D1-D3 (colour temperature, tone curve, saturation character) — free, three adds and one array increment inside a pixel loop that already runs — see Semantic visual dimensions (Phase 4)
DetectLookGroupstruePhase 4, D4 look grouping (view.lookGroups, ids k{n}) — free, O(shots²) over six floats, cheaper than the 576-float duplicate grouping already running
LookSimilarityThreshold0.88Minimum look similarity (0..1) for two shots to share a look group; lower than DuplicateSimilarityThreshold since look distance is a weighted six-vector, not a 576-float grid comparison
MaxViewLookGroups12Caps view.lookGroups
VisionOffPhase 2: Off / Optional / Required — vision-LLM shot captioning, off by default (unlike Transcription) — see Vision captioning (Phase 2)
VisionProviderIdnullExplicit override; otherwise resolved via the default Vision-capability provider
CaptionSelectionPerDuplicateGroupPerDuplicateGroup / LongestShots / EvenlySpaced — which shots get captioned; round-robins across source clips when more than one is analyzed (Phase 4)
MaxCaptionedShots50Hard cap on vision chat-completion calls across the whole step (genuinely step-wide since the Phase 4 vision hoist — see Vision captioning (Phase 2)), not a target: every shot under MinCaptionShotSeconds is excluded before this cap even applies. Default raised from 24
MinCaptionShotSeconds1.0Shots shorter than this are never selected for captioning
KeyframeMaxWidth512Max width (px) of the extracted keyframe JPEG (or contact sheet — see KeyframesPerShot) sent to the vision model; never upscaled
VisionTimeoutSeconds120Aggregate wall-clock budget for the whole captioning pass (not per-shot)
MaxCaptionChars320Caption summary field is truncated to this length
KeyframesPerShot1Phase 4: frames combined into ONE contact-sheet keyframe per captioned shot, clamped 1..3. 1 (default) is a single mid-shot still, byte-identical to the pre-Phase-4 vision path — see Semantic visual dimensions (Phase 4)
PersistKeyframesfalsePhase 4: now wired — when true, each captioned shot's keyframe JPEG is uploaded to storage under the video-analysis/{executionId}/step-{n}-keyframes/ prefix (a persist failure never costs the caption itself). When false (default), keyframes stay scratch-only and are deleted with the rest of scratch space
EmitOverlayPlacementsfalsePhase 3: derives deterministic overlay-placement candidates (view.placements) from each shot's Phase 1 region data — see Motion graphics (Phase 3)
MaxPlacementsPerShot2Top-N regions (by Suitability) offered per shot
MaxPlacements40Hard cap on placements across the whole artifact; lowest-suitability candidates dropped first
MaxTimeSlicesPerRegion3When a region's chosen time window is long enough to hold more than one distinct overlay moment, split it into up to this many non-overlapping, evenly-spaced sub-windows instead of offering every overlay the same window — see Motion graphics (Phase 3)
MaxSharpnessShots24Phase 4: step-wide ceiling on sharpness measurements when DetectSharpness is on — genuinely step-wide like MaxCaptionedShots, not per source. Costs one extra ffmpeg invocation per measured shot
AnalyzeMusicBeats / MusicBeatTrackProjectFileId / MaxBeatAnalyzedTracksfalse / null / 6Measure tempo, beat, bar, drop and outro of offered music tracks and of the track the edit is cut to, for the editor (view.musicTracks[].beat, view.music) and for the compile's BeatSync — see Cutting to the beat
OfferMusicTracksfalseEnumerates every audio/* project file as an m{n} music-track candidate (view.musicTracks) for a downstream AgentType.MusicSupervisor step — project-level, not per-source. See Background music
MaxMusicTracks20Caps view.musicTracks
OfferSfxClipsfalseEnumerates every audio/* project file as an x{n} SFX-clip candidate (view.sfxClips) for a downstream AgentType.SoundDesigner step — project-level, sharing one ListFilesAsync call with OfferMusicTracks when both are on. See Sound effects
MaxSfxClips40Caps view.sfxClips
DetectInsertRegionsfalsePhase 5: chroma-plate quad tracking for tracked screen inserts — one extra medium-res grid ffmpeg pass per source + pure C# (ChromaQuadTracker); tracks offered as r{n} ids (view.insertRegions) — see Tracked screen inserts (Phase 5)
MaxInsertPlatesPerFrame1Chroma plates tracked SIMULTANEOUSLY per frame — the multi-plate capability. 1 is the original largest-component-only behaviour exactly; clamped 1..16 — see Non-planar surfaces and multi-plate compositing
MeasureInsertCurvaturetrueMeasure each plate's silhouette (refined corners + edge curvature). The refined corners improve the FLAT composite too; false reproduces the pre-existing composite exactly — see Non-planar surfaces and multi-plate compositing
SolveInsertChromaKeytrueGrid-search each tracked plate's own colorkey parameters so the compile step can clip the insert to the plate's real per-frame silhouette — rounded bezel corners, a camera notch, a hand crossing the screen — see Silhouette matting
InsertRegionColor"green""green"/"blue"/"magenta" — matched entirely in C# channel-ratio space, NEVER an ffmpeg value; unknown values fall back to green
InsertSampleFps0 (auto)Sample rate of the dedicated tracking pass. Auto = the source's own frame rate, or the smallest exact divisor of it ≤ 30 fps that fits the frame cap and memory budget; a positive value is an explicit override (clamped 0.5..30). See Sub-pixel corners at the source's frame rate
InsertGridWidth / InsertGridHeight0 / 0 (auto)Tracking grid resolution (clamped 64..640 / 36..360). 0 derives it from the SOURCE's probed aspect (~320px long axis), so grid-space and source-space distances agree — see "The tracking grid preserves the source's aspect ratio". Set BOTH to override
MaxInsertSampleFrames3000Clamps effective tracking fps downward for long videos, exactly like MaxVisualSampleFrames
MinInsertRegionAreaRatio0.004Minimum fraction of frame area a chroma component must cover to count as a plate
MinInsertRegionSeconds1.0Tracks shorter than this are dropped
MaxInsertRegions8Cap on offered tracks across the whole artifact (longest kept)
ExpectnullOptional structural checks (MinShots, MinTranscriptSegments, MaxSilenceRatio, MinShotsWithVisuals)

VideoCompileStepConfig​

FieldDefaultNotes
Version—Config schema version
Decision—ExtractInputRef (reused verbatim from StepType.Extract) — only From = Previous or From = Step are valid here
AnalysisStepOrder—Which VideoAnalyze step's artifact to resolve ids against
AnalysisStepResultIdnullCross-execution override — resolve a prior run's artifact instead of this execution's; scoped to the same project
ModeReencodeReencode (frame-accurate select/aselect filtergraph) or StreamCopy (fast, lossless, but cuts snap to keyframes)
PrePaddingMs / PostPaddingMs80 / 120
MinSegmentMs250Spans shorter than this are dropped
MaxSegments200Above ~64 segments the filtergraph is written to a scratch file and passed via -filter_complex_script to avoid argv length limits
AllowKeyframeSnappingfalseMust be true to use Mode = StreamCopy
OutputFileName"edited.mp4"Sanitized to a safe character set with a forced extension. When left at the default and RegisterProjectFile is on, the project file is registered as "{workflow name} - run {n}.mp4" (plus (revision {k}) for a review-loop re-compile) instead — see Compile warnings for the naming rule. A custom name is kept verbatim
OutputFormatSourceSource (canvas = the largest kept source), Vertical (1080x1920), Square (1080x1080), Landscape (1920x1080). A fixed format requires Mode = Reencode. See Output format
OutputFitCropHow a clip whose aspect differs from a fixed format is fitted: Crop (scale to cover, centre crop) or Pad (scale to fit, filled per CanvasFill). Ignored for Source
TargetDurationSec0The length the customer asked for; 0 = none. A longer cut is trimmed to it (Target length); reported as length in the step output, with a target_length_missed warning when off by more than 20%
VideoCodec"libx264"Allowlisted (libx264, libx265, libvpx-vp9) — config is workflow-author-supplied, not model output, but still reaches ffmpeg argv. Whatever the codec, a re-encoded output is always 4:2:0 (yuv420p), standard range, with the MP4 index up front (+faststart), and H.265 is tagged hvc1: without an explicit pixel format the encoder followed the filtergraph to full-range or even 4:4:4 output (H.264 "High 4:4:4 Predictive" once tracked inserts composited in planar RGB), which played on a desktop and was refused by phones and WhatsApp
AudioCodec"aac"Allowlisted (aac, libmp3lame, copy)
Crf20Clamped 0..51
Preset"veryfast"Allowlisted (ultrafast … veryslow)
RegisterProjectFiletrueRegisters the compiled video as a re-editable ProjectFile row
GraphicsPlannullPhase 3: ExtractInputRef (Previous/Step only) — which step's resolved MotionGraphicsPlanOutput to apply. null = no graphics looked up. See Motion graphics (Phase 3)
EnableGraphicsfalsePhase 3: applies the resolved graphics plan during the same encode. false (default) is byte-identical to the pre-Phase-3 compile path. Requires Mode = Reencode
MaxOverlays20Cap on applied overlays; excess dropped (recorded in the graphics block)
OverlayShortMs / OverlayMediumMs / OverlayHoldMs1500 / 3000 / 6000Milliseconds an overlay stays on screen, keyed by the model's Duration word (Short/Medium/Hold)
OverlayFadeMs300Fade-in/fade-out duration at each end of an overlay's on-screen window
OverlayFontSizePct5Percent of frame height; clamped 2..12 at execution time
OverlayBoxHeightPct16Percent of FRAME height the drawn overlay box (drawbox/drawtext background, or the box a rendered-asset overlay is stretch-scaled into) occupies — a compact accent strip, not the named safe-zone band's own height. Clamped 6..40, never exceeding the band's own height
OverlayBoxWidthPct82Percent of the named band's own WIDTH the drawn overlay box occupies, centered. Clamped 30..100
OverlayFontColor"white"Allowlisted (white/black/yellow/#RRGGBB) — reaches ffmpeg's filter string, so validated like VideoCodec
OverlayBoxColor"[email protected]"Allowlisted ([email protected], [email protected], [email protected], none)
MaxOverlayTextChars / MaxOverlaySubtextChars80 / 60Sanitized-text truncation budget (OverlayTextSanitizer)
MusicPlannullBackground music: ExtractInputRef (Previous/Step only) — which step's resolved MusicPlanOutput to apply. null = no plan looked up (deterministic MusicTrackProjectFileId path, or no music, is used instead). See Background music
MusicTrackProjectFileIdnullA specific audio/* project file to use as the music track — the deterministic path (no agent required), and also the fallback when MusicPlan is unresolvable/invalid or names an unoffered track id
BeatSyncOffBeat / Bar: snap every cut to the music's beat or bar by trimming span tails, and cut on the drop — see Cutting to the beat. Off is byte-identical
EnableMusicfalseApplies the resolved music track during the same encode. false (default) is byte-identical to the pre-music compile path. Requires Mode = Reencode and AudioCodec != "copy"
MusicDuckingSpeechEnvelopeOff / SpeechEnvelope — SpeechEnvelope lifts the music during non-speech windows via a deterministic keyframed volume envelope; Off is a constant ducked bed throughout
MusicFitPolicyLoopToFitLoopToFit (loops via -stream_loop -1 to fill the whole edit, then trims to its exact length) / PlayOnce (plays once, trimmed to its own length if shorter than the edit)
MusicFadeInMs / MusicFadeOutMs1500 / 2500Fade duration at the start/end of the music track's own play window
MusicBedQuietDb / MusicBedBalancedDb / MusicBedFeatureDb-26 / -20 / -14Bed level (dBFS), keyed by the model's Intensity word (Quiet/Balanced/Feature). Clamped [-40, -6]. When the output has no dialogue and no narration (silent footage, or MuteSourceAudio) the music is the whole soundtrack and plays at -1 dB regardless, through a −2 dBFS peak limiter (alimiter, lookahead compensated) because a mastered track already decodes over full scale (music.bedBasis: "sole_soundtrack")
MuteSourceAudiofalseDrop the clips' own sound from the program: only music, narration and sound effects are heard. For generated clips that carry their own (often unwanted) audio. audio.reason reads source_audio_muted when it removed real audio
MusicDuckLightDb / MusicDuckNormalDb / MusicDuckHeavyDb-6 / -11 / -18Attenuation (dB, below the bed) applied while dialogue is present, keyed by the model's Ducking word. Clamped [-30, 0]
MusicDuckRampMs400Linear gain ramp (ms) INSIDE each lift window — the music is never above the ducked level exactly at a window boundary
MinMusicLiftWindowMs1200Non-speech windows shorter than this (and shorter than twice the ramp) are never lifted at all
MaxMusicLiftWindows12Caps the volume-envelope expression's length; excess windows dropped, longest first, then re-sorted chronologically
MusicLiftMergeMs400Lift windows closer together than this are merged into one
TransitionPolicyOff*VideoTransitionPolicy — gates the deterministic seam-transition system; Off is byte-identical to the pre-transition compile path. The video-derush-edit* templates opt in with "Auto". See Seam transitions and the program envelope
AudioSeamRampMs—Duration of the audio-only declick ramp treatment at a cut seam
SoftCutMs—Duration of a soft-cut (brief cross-blend, shorter than a full dissolve) treatment
DissolveMs—Duration of a full crossfade-dissolve treatment
DipToBlackMs—Duration of a dip-to-black treatment (fades to black, then to the next shot)
DipCutMs—Duration of a dip-cut treatment (a very brief dip, shorter than a full dip-to-black)
MaxTransitionMs—Hard cap on any single transition's duration, regardless of treatment
MaxTransitionRatioPct—Caps a transition's duration as a percentage of the SHORTER of its two neighbouring segments, so a transition can never eat a meaningful fraction of either one
MaxTransitionSegments—Above this many segments in the compile, transitions are skipped entirely (every seam falls back to a hard cut) — a filtergraph-buffering guardrail, not a quality knob; see Seam transitions and the program envelope
SectionBreakGapMs—Minimum silence-gap duration at a seam for the rule table to treat it as a section break rather than an ordinary mid-sentence cut
ProgramFadeInMs / ProgramFadeOutMs—Video fade-in/fade-out duration at the very start/end of the whole compiled program (distinct from any inter-cut transition)
ProgramAudioFadeInMs / ProgramAudioFadeOutMs—Audio fade-in/fade-out duration at the very start/end of the whole compiled program, tracked independently of the video program fade
EnableInsertsfalsePhase 5: composites the plan's chosen screen inserts (corner-pinned via ffmpeg's per-frame-animated perspective filter) during the same encode. false (default) is byte-identical to the pre-inserts compile path. Requires Mode = Reencode (INSERTS_REQUIRE_REENCODE); inserts are read from the SAME GraphicsPlan-referenced MotionGraphicsPlanOutput (EnableGraphics itself need not be on) — see Tracked screen inserts (Phase 5)
InsertSurfaceAutoWhether a plate may be composited with a piecewise-projective MESH warp instead of one corner pin. Auto decides per plate from measured curvature; Planar never meshes; Mesh always tries — see Non-planar surfaces and multi-plate compositing
MinInsertCurvature0.01Measured curvature a plate must exceed under Auto to earn a mesh, as a fraction of its own size
InsertMeshTolerancePx0.6Target worst-case chord error per mesh cell — this is what sets the cell count
MaxInsertMeshCells24Ceiling on rows * cols for one insert's mesh
EnableInsertMattetrueClip each composited insert to the plate's own per-frame silhouette, so anything passing in FRONT of the plate occludes the insert instead of being painted over. A plate with no solved key composites with the plain quad mask regardless — see Silhouette matting
InsertReflectionstrueScreen-blend the plate's own glare and reflection streaks over a matted insert, so the new screen content carries the reflections the filmed glass had. A clean plate contributes nothing — see Reflections
MinInsertConfidence0.5Tracks below this confidence are dropped (confidence_below_threshold) rather than composited badly
MaxInserts3Cap on applied inserts; excess dropped in plan order
MaxInsertExprKeyframes1500Per-insert cap on corner keyframes baked into the perspective expressions (uniform downsample; clamped 2..5000). Past 96 keyframes the series is written as a log-depth balanced tree, which is what lets a track sampled at the source's frame rate reach the encode undecimated
InsertOverscan0.02Fractional outward expansion of the tracked quad about its centroid, hiding the plate's edge fringe under the insert (clamped 0..0.1)
SfxPlannullSound effects: ExtractInputRef (Previous/Step only) — which step's resolved SfxPlanOutput to apply. null = no SFX looked up. Unlike music, deliberately NO deterministic no-agent fallback field — see Sound effects
EnableSfxfalseMixes the resolved SFX cues into the output audio during the same encode. false (default) is byte-identical to the pre-SFX compile path. Requires Mode = Reencode and AudioCodec != "copy" (SFX_REQUIRES_REENCODE/SFX_REQUIRES_AUDIO_REENCODE)
MaxSfxCues8Cap on applied cues; excess dropped in plan order (max_cues_exceeded)
MaxSfxCueSeconds4.0Hard cap on any single cue's play window — a long file misused as a cue is trimmed, never a de-facto bed
SfxSubtleDb / SfxNormalDb / SfxStrongDb-18 / -12 / -6Cue gain (dBFS), keyed by the model's Volume word. Clamped [-40, 0]
SfxLeadMs / SfxLagMs150 / 150How far a Timing: "Lead"/"Lag" cue fires before/after its anchor's output-timeline start moment
SfxFadeOutMs120Declick fade-out at the end of each cue's play window, capped at half the window
ColorGradePlannullColor grading: ExtractInputRef (Previous/Step only) — which step's resolved ColorGradePlanOutput to apply (a solo Agent(Colorist) step or a ColorGradeRoom step; both emit the same shape). null = no grade looked up. See Color grading
EnableColorGradefalseApplies the resolved colour grade (a first-party eq/colorbalance/colorlevels/hue chain keyed by the plan's enum words — ColorGradeFilterBuilder) during the same encode, before overlays/inserts. false (default) is byte-identical to the pre-grade compile path. Requires Mode = Reencode (COLOR_GRADE_REQUIRES_REENCODE); every plan-level failure degrades to "no grade applied"
ExpectnullOptional structural checks (MinOutputSeconds, MaxOutputSeconds, MinRetainedRatio default 0.15, MaxRetainedRatio)
KeepOrderSourceOrderAsListed plays every kept span in the editor's listed order, even within one clip — see Playing spans in the editor's order. SourceOrder is byte-identical to before

* TransitionPolicy and the twelve fields above it (AudioSeamRampMs through ProgramAudioFadeOutMs) belong to a sibling in-flight change adding the deterministic seam-transition system to VideoCompileStepExecutor/VideoCompileStepConfig; as of this doc edit they are not yet present on the shipped VideoCompileStepConfig type in this worktree, only referenced by name in the video-derush-edit* templates' seeded config JSON and in this section, so their defaults are intentionally left blank above pending that merge.

The video-derush-edit template​

An opt-in workflow template (AutoCreateOnProject: false, seeded in ReelBolt.Shared/Workflows/WorkflowTemplateCatalog.cs) demonstrating the full pipeline: VideoAnalyze (Source: PreviousStepOutput) → Agent(VideoStoryEditor) → VideoCompile (Decision: Previous, AnalysisStepOrder: 1). Its literal seeded JSON is deserialization-tested against the real config types in WorkflowTemplateCatalogConfigDeserializationTests.cs, so a future field-name drift between the template and the config records it targets fails CI rather than a live workflow run.

The video-derush-edit-graphics template​

A fourth opt-in template (AutoCreateOnProject: false), extending video-derush-edit with Phase 3 motion graphics end to end: VideoAnalyze (Source: ProjectFile, EmitOverlayPlacements: true) → Agent(VideoStoryEditor) → Agent(MotionGraphicsPlanner) → VideoCompile (Decision: Step 2, AnalysisStepOrder: 1, EnableGraphics: true, GraphicsPlan: Step 3). Decision/GraphicsPlan reference their source steps explicitly by StepOrder rather than Previous, since Previous relative to the compile step would resolve to the MotionGraphicsPlanner step's output, not the story editor's decision. Deserialization-tested the same way as video-derush-edit.

The video-derush-edit-music template​

A fifth opt-in template (AutoCreateOnProject: false), extending video-derush-edit with background music instead of graphics: VideoAnalyze (Source: ProjectFile, OfferMusicTracks: true) → Agent(VideoStoryEditor) → Agent(MusicSupervisor) → VideoCompile (Decision: Step 2, AnalysisStepOrder: 1, EnableMusic: true, MusicPlan: Step 3). Same explicit-StepOrder rationale as video-derush-edit-graphics above (Previous relative to the compile step would resolve to the MusicSupervisor step's own output, not the story editor's decision). See Background music. Deserialization-tested the same way as the other two templates.


Multiple source clips​

VideoAnalyzeStepConfig.Sources (plural — IReadOnlyList<VideoSourceRef>) lets one VideoAnalyze step analyze several source clips (e.g. multiple takes, camera angles, or B-roll of the same scene) into ONE merged artifact, which a single VideoStoryEditor decision and a single VideoCompile step can then cut across. It is a strict superset of the original single-clip behavior: Sources null/empty (the default) is treated as a one-element [Source] list, so every existing persisted/template config — which only ever set the singular Source — keeps deserializing and behaving byte-identically. A one-element Sources list behaves identically to the equivalent single-Source config too; there is no separate "N=1" code path anywhere in this addition.

Analysis: independent per-clip passes, merged into one global id space​

VideoAnalyzeStepExecutor analyzes each clip independently via the exact same deterministic per-source pipeline it always ran (silence/shot detection, transcription, Phase 1/2/3/4 analysis), processed sequentially, never in parallel, one clip's local file at a time. The pre-decode guardrails (MaxDurationSeconds, MaxInputBytes) are enforced per source clip, not summed across all of them.

Each clip's own ids restart at s0/g0/t0/w0/p0/d0. VideoAnalyzeStepExecutor.OffsetId then remaps every local id into ONE globally-unique id space via a running per-id-kind offset (assigned contiguously across every source file, in analysis order), and every shot/silence gap/segment/word/placement/duplicate-group is tagged with the SourceIndex of the clip it came from. The bounded view surfaces this as a "src" index on every offered item (e.g. "src": 0), so the VideoStoryEditorAgent prompt can tell the model which clip each id belongs to and instruct it to freely alternate between clips across successive Keep spans — picking whichever clip has the best material for each moment is the whole point of offering more than one. The one hard rule: a single Keep span's FromId and ToId must both come from the SAME clip, since a span is a contiguous run within one physical file, never a bridge across two files — cross-clip edits are expressed as a SEQUENCE of single-clip Keep spans instead.

Per-source provenance (VideoAnalysisProvenance) is aggregated into one artifact-level record via AggregateProvenance: applied/degraded flags become true if ANY source applied/degraded that stage — an artifact-wide OR, not a per-source breakdown. Full per-source provenance detail is deliberately out of scope for this addition; a single source passes through unchanged (Count == 1 short-circuits).

VideoAnalysisArtifact.Sources records one VideoAnalysisSourceInfo per analyzed clip, in source-index order — the storage key VideoCompileStepExecutor must download to physically cut from that clip, and the per-clip VideoAnalysisMedia (duration/fps/dimensions) every clip-aware computation (frame quantization, padding clamps, graphics geometry) must use instead of a single artifact-wide Media. null only for a true legacy artifact produced before this field existed — the one case VideoCompileStepExecutor still re-derives the source storage key the old way, by walking the VideoAnalyze step's own config. Every artifact produced by the current executor populates Sources with at least one entry, even for a single source, so the top-level Media and Sources[0].Media always agree for that case.

Compile: which ids can pair, and how the cut is resolved​

VideoCompileStepExecutor resolves each Keep span's FromId/ToId to [Start, End) times plus the SourceIndex recorded against that id — never trusted from the model, since there is no source field on VideoEditKeepSpan for it to get wrong in the first place:

  • Spans across clips — a single span whose FromId and ToId resolve to different SourceIndex values is SPLIT, when its ids run forward in the offered order, into one span per clip (SplitAcrossSources: consecutive same-source runs, the first keeping the span's own FromId so its transition still applies), reported as spanNormalization.splitAcrossSources. Every still image is its own source, so "the last beat of one photo to the first of the next" is an ordinary thing for an editor to write, and failing the workflow on it was worse than the obvious reading. A span whose ids run BACKWARDS across clips still fails with MIXED_SOURCE_SPAN, before any normalization runs.
  • Ordering/overlap/coalescing only ever compares a span against the immediately PRECEDING span in list order; when that neighbor belongs to a DIFFERENT source clip, there is no shared timeline to be "out of order" or "overlapping" on, so the check (and coalescing) is simply skipped at that boundary. A single-source config's spans are always same-source neighbors, so this reduces to exactly the original single-timeline behavior.
  • Padding/clamping clamps each span against ITS OWN source clip's duration (GetSourceMedia), never a single artifact-wide duration.
  • Frame-quantization uses each span's OWN source clip's exact rational fps, never a single artifact-wide fps.
  • Retained ratio (Expect.MinRetainedRatio) sums only the DISTINCT clips actually referenced by the resolved cut list, not every clip the step merely analyzed — for a single source this sum has exactly one term, so it is byte-identical to before this addition.
  • Only the DISTINCT source clips actually referenced by the resolved cut list are downloaded — never every clip the VideoAnalyze step analyzed.
  • The "canonical" clip every multi-source encode normalizes toward (scale/pad/fps for video, sample rate/channel layout for audio) is the LARGEST kept source by pixel area, ties going to the clip kept first — still deterministic, independent of clip count or offered-id ordering. It used to be the first kept span's clip, which shrank a narrated edit's 1920x1080 b-roll to the 1344x768 of the generated clip that happened to open it. Every geometry computation (overlays, inserts, censor masks) reads the canvas through the same per-clip fit the encode applies; see Output format.

MULTI_SOURCE_REQUIRES_REENCODE — when the resolved cut list references more than one distinct source clip, Mode = StreamCopy is refused as a hard, pre-encode config error (the same discipline GRAPHICS_REQUIRE_REENCODE already established for a different Reencode-only combination): losslessly concatenating independently-encoded files has no correctness-preserving stream-copy equivalent, since ffmpeg's concat filter/demuxer both require matching codec parameters across inputs that separately-encoded source files are not guaranteed to share, and normalizing them first is itself a re-encode.

Why multi-source encoding is a separate method​

EncodeReencodeMultiSourceAsync is a deliberately separate method from the original EncodeReencodeAsync — which stays completely untouched, and is still used for every single-source compile, so a single-source (or single-clip-in-practice) compile's ffmpeg argv/behavior stays byte-identical to before this addition. It is not a generalization of the single-source method because ffmpeg's select filter always emits one input's own matched ranges in THAT INPUT'S OWN chronological order — it cannot express a Keep-span order that jumps between clips arbitrarily. Only the concat filter, fed one small pre-trimmed clip PER SPAN in the exact order they should play, can. Each span becomes its own trim/atrim branch off the correct ffmpeg input index for that span's own source clip, normalized to the canonical frame size/rate/audio format concat requires, then concatenated in Keep order.

Audio-less source clips​

Real B-roll/stock footage routinely ships with no audio stream at all. VideoCompileStepExecutor probes every distinct referenced source once (cheap) before encoding and threads the result through as a per-source sourceHasAudioByIndex map:

  • If NOT ONE referenced clip has an audio stream and nothing is mixed in, the whole compiled output drops audio entirely — there is nothing to preserve, so an all-silent track would add nothing. When narration or sound effects ARE mixed in, a silent base is synthesized instead; see Program audio on silent footage.
  • Otherwise (a mix of audio-having and audio-less clips), concat's own stream-count contract (every concatenated segment must carry the same a= count) is satisfied per span: a span whose own source clip has audio gets its real atrim branch; a span whose source clip has no audio instead gets a synthesized, matching-duration silence branch (ffmpeg's anullsrc source filter, already natively 48kHz/stereo, so it needs no extra -i input or aformat). Every span with real dialogue keeps its real dialogue — only the audio-less span(s) carry synthesized silence. The step's output JSON records which resolved-span indices got synthesized silence (audio.syntheticSilenceSegmentCount/syntheticSilenceSegmentIndices), so this is never a silent surprise the way an unreported drop would be.
  • Whether the FINAL OUTPUT has any dialogue audio at all (hasDialogueAudioInOutput — isMultiSource ? anySourceHasAudio : sourceHasAudio) also feeds Background music's ducking decision: there is no point planning silence-gap ducking windows against dialogue that will not exist in the output.

Playing spans in the editor's order​

By default a compile plays the spans of one source clip in that clip's own order, whatever order the editor listed them in: within each run of consecutive same-clip spans they are sorted by start time and overlaps merged (outputSummary.spanNormalization reports what moved). Spans of different clips already played in the listed order. So "open on the strongest shot" was impossible when that shot was filmed last — story order was source order.

VideoCompileStepConfig.KeepOrder (VideoKeepOrder, append-only, persisted by name) changes that:

ValueBehaviour
SourceOrder (default)Exactly the behaviour above — byte-identical
AsListedEvery span plays in the editor's listed order, within one clip too

Under AsListed nothing is sorted. Two rules keep the cut well-formed:

  • No moment plays twice. RemoveEarlierCoverage removes from each span every part an EARLIER span (in list order) of the same clip already covers. A span split in two by an earlier one keeps both pieces (in source order); one covered entirely disappears. It runs again after padding, so a padded boundary never replays a few frames of an earlier span.
  • Only forward continuations merge. CoalesceInListOrder merges a span into the one before it only when it starts exactly where that one ends in the same clip. A span that jumps back is a deliberate reorder and stays its own span.

select/aselect (the single-clip encode) can only play a file forwards, so when any span of a clip plays before an earlier part of the same clip (CountOutOfSourceOrder > 0) the compile takes the segmented trim + concat encode, even for one source. With Mode = StreamCopy that is a hard config error, REORDER_REQUIRES_REENCODE. An AsListed decision that happens to be in source order produces exactly the default cut. The seam planner treats a backward jump within one clip like a cut between clips (no "removed gap" is measured for it).

The output summary carries, only under AsListed:

"keepOrder": { "mode": "AsListed", "outOfSourceOrderSpans": 1, "overlapTrimmedSpans": 0 }

The decision schema is unchanged — VideoEditDecisionOutput was always an ordered list, and its no-timestamp invariant tests are untouched. The VideoStoryEditor prompt (both copies, the agent's DefaultPrompt and DatabaseSeeder) gained a separate section, "Playing spans out of source order (only when the brief asks)": list spans in the order they should be seen when the brief asks for a different order, never list a moment twice, and keep source order otherwise. On a compile left at SourceOrder a reordered list is simply put back in source order, so the prompt can never break a cut.


Still images​

An image is a source like any clip: a ProjectFile (or step output) whose key ends in .png, .jpg, .jpeg, .webp, .bmp, .tif/.tiff or .avif — or, with no extension, whose probe reports an image codec and no duration — is turned into video once, in VideoAnalyzeStepExecutor.RenderStillAsync, and the clip is recorded as the source exactly like an author-trimmed window, so the editor, captions, music, grading and the compile all work on it unchanged.

The clip (StillImageClip) is the photo at its own aspect (long side at most 1920, even sides), upscaled 2x before zoompan so the moves do not judder, with three MOVES back to back — a push in (8%), a pull back out, and a pan along the long side at 1.12x — each two 1.5 s beats, 9 s in all, no audio. Each beat is declared as an author scene (CompositionSceneManifest, IsComposition: false), so the analysis offers six shots per photo, named push in (1/2) … pan right (2/2) (pan down for a tall photo). The FIRST beat's caption leads with what the photo shows — ONE vision caption of the image when Vision is on and a provider resolves, "A still photo." otherwise, capped at 200 characters — and every other beat says "Same photo as this source's first shot" and its move: six copies of one caption per photo pushed a 13-source view over its budget, and the view's only remedy (detail down to None) stripped every caption, move names included. The editor therefore chooses the motion by which shots it keeps and its length by keeping one beat (1.5 s) or both (3 s) — all by id. The VideoStoryEditor prompt (both copies) has a "Photos" section: the six shots are alternatives, so keep ONE move per photo (a first run kept all six beats of every photo, 74 s for a 25 s brief). meta.stillImages lists the still sources. InSec/OutSec on an image are ignored; a render failure fails the step with STILL_IMAGE_FAILED.

Using only part of a source​

A workflow author can use only a window of a source clip: VideoSourceRef gained InSec and OutSec (seconds of the original file; either may be null for "from the start"/"to the end"; appended to the positional record, so every stored payload deserializes unchanged and both null is byte-identical). The web clip list shows Use from / Use until (seconds) on every clip.

VideoAnalyzeStepExecutor.TrimSourceAsync handles it once per source, right after the download:

  1. The window is validated against the file's probed length (InSec >= 0, at least 0.1 s long, the start inside the file; an OutSec past the end is clamped to it). Anything else fails the step with SOURCE_TRIM_INVALID before any decode.
  2. The window is cut out frame-accurately (input seek + re-encode, BuildTrimArgs: libx264 -crf 16, the source's own frame rate, its audio when it has any) and uploaded as the step's own artifact (video-analysis/{executionId}/step-{n}-src{i}-window.mp4).
  3. Everything else — probe, silence and shot detection, transcription, the visual grid, insert and object tracking — runs on the WINDOW, and the artifact records the window as the source (VideoAnalysisSourceInfo.StorageKey/Media), plus TrimInSec/TrimOutSec for the record.

So every id, time and track is relative to the window's own start, and no later step needs to know a trim happened: VideoCompileStepExecutor downloads the recorded window and maps ids through it exactly as for any source. Author-declared scenes (a sidecar next to a render) describe the whole piece, so a trimmed source detects its shots instead. The envelope's meta.sourceWindows ([{ src, inSec, outSec }]) lists the windows, present only when a source was trimmed. Trimming needs the video tool runner (always registered in the engine; a construction without it fails with SOURCE_TRIM_UNAVAILABLE), and SOURCE_TRIM_FAILED reports an ffmpeg failure.


Keeping part of a long shot​

An editor that may only name ids could keep a long shot whole or drop it whole — a great shot with a bad second at its end had to go. VideoAnalyzeStepConfig.MaxShotSec (default 0, off) splits every detected shot longer than that (values under 1 s act as 1 s) into n = ceil(length / MaxShotSec) sub-shots (ShotSplitter.Split). Each split starts at the even division and moves, within 30% of a part either way, to:

  1. the stillest moment there — the mean absolute luma change between consecutive frames of the Phase 1 visual grid (ShotSplitter.MotionCurve, smoothed over three samples). The grid is sampled once, before the split, and reused by the visual analysis, so this costs no extra decode; then
  2. the middle of the nearest pause (a detected silence) when no motion was measured; then
  3. the even point. A move that would leave a part shorter than 0.5 s falls back to the even point.

Sub-shots are ordinary shots with consecutive s{n} ids — the editor keeps part of a long shot by keeping only some of them, still never naming a time — and carry Part ("2/3") in the artifact and part in the view (absent on every unsplit shot). Splitting happens before ids are assigned and before audio levels, captions and duplicate grouping, so each part gets its own measurements. Author-declared scenes are never split. The VideoStoryEditor prompt (both copies) gained a short "Parts of a long shot" section.


Output format​

VideoCompileStepConfig.OutputFormat picks the canvas the program is encoded at, and OutputFit how each clip is fitted into it. Source (default) keeps the clips' own size: a single clip is encoded at its own resolution exactly as before, and several clips share the canvas of the largest kept source (pixel area, ties to the clip kept first) — never merely the first clip's, which used to downscale 1080p b-roll to whatever generated clip opened the edit. The fixed formats are Vertical 1080x1920 (Reels, Shorts, TikTok), Square 1080x1080, Landscape 1920x1080 and Portrait 1080x1350 (4:5 feed posts). A run can override OutputFormat on every compile step at once — the quick format switch, which also turns the default Crop fit into Subject for a portrait or square target — and the analyze view then tells agents each clip's orientation and the target; see editing-styles.md.

OutputFitWhat a clip of a different aspect gets
Crop (default)scale=W:H:force_original_aspect_ratio=increase,crop=W:H — fills the frame; a 16:9 shot in 9:16 keeps its middle third
Padscale=W:H:force_original_aspect_ratio=decrease,pad=W:H — the whole shot, filled per CanvasFill (black bars, or a blurred copy of the shot)
Subjectthe Crop scale, with the crop window placed on the clip's subject instead of its middle — see Reframing to the subject

Both encode paths fit. The single-source select path appends the crop/pad suffix to its cut stage, before the grade, inserts, overlays and captions, so everything painted afterwards works in canvas pixels. A single clip at Pad + CanvasFill = Blur needs a split/blur/overlay graph per clip, so that one combination routes to the segmented path, which builds it. On the segmented path a fixed format normalizes every span (fitAllSourcesToCanvas), even when there is only one clip.

Geometry follows the picture through one mapping, SourceCanvasFit.For(source, canvas, cover): the contain-fit arithmetic pad uses, or with cover the scale-to-cover arithmetic crop uses, whose offsets go negative because the picture overflows the canvas. Tracked-insert corners, censor masks and overlay placement bands (MapRectThroughFit, clamped to the frame) are all mapped through it; a placement band that a crop removes entirely is dropped as placement_cropped_out.

A fixed format with Mode = StreamCopy is the hard config error OUTPUT_FORMAT_REQUIRES_REENCODE — resizing has no stream-copy equivalent. Every compile reports what it encoded:

"output": { "format": "Vertical", "width": 1080, "height": 1920, "fit": "Crop" },
"length": { "targetSec": 20, "actualSec": 10.0, "met": false }

output.fit is the configured fit at a fixed format, Pad for several clips on a source-sized canvas, and None for a single clip at Source. length is present only when TargetDurationSec > 0; met is false when the real length is more than 20% away from the target. The compile never invents footage to meet a target — a short source makes a short video, and the customer is told so (target_length_missed, see Compile warnings).

Target length​

A cut that runs LONG is brought down to TargetDurationSec before the encode (and before cutting to the beat, which only shortens further) by TargetLengthPlanner: span tails are shortened under one common ceiling, longest spans first, so every kept span keeps its opening and short beats are untouched. No span goes below the narration anchored in it (the same floors beat sync uses) or 1 s, and the LAST span — the end card or closing beat — is never touched. When those floors make the target unreachable the cut comes as close as they allow. length.fitted reports editedSec (the editor's length), trimmedSec and trimmedSpans; it is absent when nothing was trimmed. A POV teaser asked to run 15 s came out at 25 s while the length was only reported.

Reframing to the subject​

A centre crop is wrong whenever the subject is not in the middle: reframing a 16:9 shot to 9:16 keeps only x ∈ [0.34, 0.66], and a phone held at a third of the frame is cut out of its own shot. OutputFit = Subject keeps the Crop scale but places each clip's window on what the clip is about. The position is chosen per source clip, from what the analysis measured inside the parts of that clip the edit KEEPS (CropFocusResolver), in this order:

  1. A tracked screen — every insert/object-track keyframe of that source inside a kept window (basis: "tracked"). Measured every frame and free, and on a tracked-insert reel it is exactly the thing the shot is about.
  2. A located subject — VideoAnalyzeStepConfig.LocateSubjects asks the Vision provider for one box per shot on the shot's keyframe (the same IObjectLocator object tracking uses: integer 0..1000 coordinates, robust JSON repair, first box wins), then re-asks on a zoomed crop around it exactly as the object tracker refines its seeds — a full-frame box is loose (a laptop came back 0.39 of the frame wide around a 0.29-wide screen), and a loose box is what tips a clip into padding. Two calls per shot; a failed refinement keeps the coarse box. Stored as VideoAnalysisShot.Subject and weighted by how long each shot is kept (basis: "subject"). SubjectLabel names what to look for ("the phone", "the presenter"); null asks for the main subject. MaxSubjectShots (40) caps the calls, and the step's VisionTimeoutSeconds bounds them. A failed or empty lookup leaves the shot without a subject; it never fails the step. The box is in the artifact only — never in the view, never offered to an agent, and it can only ever move a crop.
  3. Otherwise the centre (basis: "centre"), which is byte-identical to Crop.

The window is static per clip. Along each axis it takes the position that keeps the most weighted subject inside it, and among equally good positions the one nearest the weighted centre: one subject that fits is centred in the window; two subjects too far apart to both fit leave the one on screen longer whole, rather than halving both (a weighted centre) or showing neither (a union centre).

When that window would keep less than 87% (CropFocusResolver.MinKeptFraction) of the clip's MAIN object — the track or located subject on screen longest — no crop can honour "keep the subject in shot", so that one clip is shown whole instead: padded per CanvasFill, exactly like OutputFit = Pad (CropFocus.Pad, reported as "fit": "Pad" on its output.focus entry). An end card whose headline is wider than a 9:16 window, or a monitor filling a 16:9 shot, keeps every word and edge; a phone or a laptop screen that fits is cropped. Secondary subjects never trigger it — a second box too far away to share the window is simply left out. A grid-saliency estimate (detail + motion on the 32x18 analysis grid) was measured first and rejected: it locked onto sharp background clutter, placing the woman's phone shot at 0.56 against a true 0.33.

The crop is crop=W:H:'max(0,min(iw-ow,fx*iw-ow/2))':'max(0,min(ih-oh,fy*ih-oh/2))' after the cover scale, on both encode paths (CoverCropFilter), and SourceCanvasFit.For(..., focus) mirrors the same clamp so tracked inserts, censor masks and overlay bands stay registered to the picture (verified pixel-identical against an explicit crop=W:H:x:0). A seam-bridge clip is always centred. The compile reports where each clip's window went:

"output": { "format": "Vertical", "fit": "Subject",
"focus": [ { "source": 0, "x": 0.33, "y": 0.5, "basis": "subject" } ] }

Program loudness​

VideoCompileStepConfig.LoudnessTarget (ProgramLoudness: Off default and byte-identical, Social −14 LUFS / −1.5 dBTP, Web −16 LUFS / −1.5 dBTP, Broadcast −23 LUFS / −2 dBTP — EBU R128) brings the finished program to the level the platform it is posted on plays back at. A narrated launch reel measured −22.5 LUFS, about 8 dB quieter than everything around it in a feed.

It runs after the encode, in two ffmpeg passes over the encoded file (LoudnessNormalizer): a loudnorm measuring pass, then an applying pass fed every measured value with linear=true — one gain for the whole program, so the balance between narration, music and effects is unchanged; when that gain would breach the true-peak ceiling, loudnorm itself switches to its dynamic mode. Only the audio is re-encoded (AAC 192k, 48 kHz); the video stream is copied, so the picture is bit-identical. The ceilings sit at −1.5 dBTP rather than −1 because the AAC encode overshoots: the same reel normalized to a −1 ceiling measured −0.6 dBTP afterwards. A peak limiter (alimiter, half a dB under the ceiling, lookahead compensated) follows loudnorm: when one linear gain cannot reach the target without breaching the ceiling, loudnorm falls back to its dynamic mode, which does not hold the ceiling — a narrated story came out at +0.5 dBTP without it and −1.8 dBTP with it (−14.3 LUFS).

The result is then measured again: the AAC encode after the limiter can still overshoot on hard transients (a hip-hop reel measured −0.14 dBTP against −1.5). When the measured true peak is more than 0.2 dB over the ceiling, the pass is redone once from the original encode with the limiter lowered by the overshoot plus 0.3 dB — one AAC generation, loudness unchanged (that reel: −3.4 dBTP, −14.3 LUFS). The node reports outputTruePeak and, when the redo ran, peakCorrectedDb.

It never fails the compile. The loudness node reports the target, the measured loudness and true peak, the gain applied, or why nothing was done (no_audio_in_program, audio_stream_copied, not_measurable for silence, measure_failed, apply_failed, normalize_failed).

"loudness": { "target": "Social", "targetLufs": -14, "targetTruePeak": -1.5, "applied": true,
"measuredLufs": -22.8, "measuredTruePeak": -3.8, "gainDb": 8.8 }

Program audio on silent footage​

Stock and generated b-roll routinely has no audio stream. Before this fix a compile over such footage with narration (or sound effects) and no music dropped audio altogether: the step reported voiceover.applied: true, captions were burned in, and the MP4 had no sound track, because the narration had nothing to be mixed into.

Now, whenever narration or sound effects are placed and neither the source nor a music bed provides a base, both encode paths synthesize one:

  • Select path (one clip). The base is anullsrc=r=48000:cl=stereo:d={program length}[abase] (BuildSilentBaseBranch) — the same 48 kHz stereo every mix branch is normalized to, ending at exactly the program's length, so each amix ... duration=first stage stays pinned to the picture.
  • Segmented path (several clips, transitions, generated clips). Every span gets the matching-duration anullsrc branch an audio-less span already gets in a mixed edit, so the concat, crossfades and cold-open/end-card concat all carry an audio stream end to end.

With music, the music branch is the base exactly as before, and a compile with nothing to mix still has no audio track; both filtergraphs are byte-identical to before the fix. The step output's audio node keeps applied/reason describing the source dialogue and adds silentBase: true, programAudio: true when a base was synthesized. Sound effects no longer need dialogue or music to be heard (the old no_base_audio drop is gone).

Compile warnings​

A render can complete and still not be what the customer asked for. The compile reports those cases in a top-level warnings array (the shared C1 shape the execution page and the renders tab display), present only when non-empty so a clean compile's output keeps its exact shape:

"warnings": [ { "code": "narration_mostly_dropped",
"message": "Only 1 of 8 narration lines made it into the video, so most of the narration is missing. ..." } ]
CodeWhen
narration_droppedSome declared narration lines (voiceover.lineCount - appliedLineCount) are not in the render, but at most half
narration_mostly_droppedMore than half are missing. The ReviewLoop also caps the review score at 4 for this shape (ReviewLoopStepExecutor.ApplyNarrationMostlyDroppedCap), so a review loop iterates instead of passing a mostly mute narrated video — see voiceover.md
no_program_audioNarration, music or effects were placed but the program has no audio track. Defensive — the silent base makes it unreachable
music_no_tracksEnableMusic is on, no fixed MusicTrackProjectFileId, and the analyze step offered no music candidates — the project has no audio files to choose from; the message tells the customer to upload one
sfx_no_clipsEnableSfx is on, nothing was applied, and the analyze step offered no SFX candidates
target_length_missedTargetDurationSec > 0 and the real length is more than 20% away from it

Output file names. With RegisterProjectFile on and OutputFileName left at edited.mp4, the render is registered as "{workflow name} - run {n}.mp4", where n counts this workflow's executions started up to this one, plus (revision {k}) when a review loop re-compiled. The name is reduced to ASCII letters, digits, spaces and ( ) - _ . , (accents are folded, anything else becomes a space — it travels into object-store metadata, which must be ASCII) and capped at 120 characters. When no workflow name can be found the render keeps edited.mp4; a custom OutputFileName is always kept. Storage keys are unchanged (they are id-based).


One registered render per run​

With RegisterProjectFile on, every compile used to add another project-file row — every review-loop pass, every retry, and every run of a workflow whose OutputFileName is fixed. One customer project ended up with ~25 identical "20s narrated timeline teaser - run 2.mp4" entries in its Files tab.

After registering its render, the compile step now removes the earlier rows THIS workflow step registered that the new render supersedes (VideoCompileStepExecutor.RetireSupersededRegistrationsAsync):

  • every earlier render of the same step in the same execution (review-loop passes, retries), whatever it was called — the Files tab keeps the run's latest render;
  • earlier renders of the same step from other executions only when they were registered under the very same file name (a fixed custom OutputFileName). Renders named "… - run {n}.mp4" are distinct per run and are kept: that is the per-run versioning.

A row counts as "this step's render" only when a WorkflowStepResult of this same WorkflowStep names its storage key as OutputStorageKey — never by name alone — so a file a user uploaded, or another step's render, is never touched. Only the rows are removed; the stored videos are not, and every earlier render stays reachable from its own run's outputs (the Renders tab and the run page read step results, not project files). The cleanup is best-effort: a database error leaves the Files list as it was and the compile still succeeds.

When it removed anything, the output summary carries:

"registration": { "fileName": "Spring teaser - run 2.mp4", "projectFileId": "…", "replacedEarlierRenders": 3 }

(absent otherwise, so a first registration's output keeps its exact shape). A workflow that pointed a later step at an earlier render's project-file id by hand will no longer find it; point it at the latest render, or at the step's output (StepOutput).


Takes: the latest render first​

Every run, review-loop pass and retry renders again, and the Renders tab and the Files view used to list every one of those renders side by side. They are now treated as takes of one video.

What a take is. Every render (a WorkflowStepResult with an OutputStorageKey under the project's outputFiles/) of the same workflow step is a take of that step's video. The group key is {workflowDefinitionId:N}:{stepOrder}:{stepType} (so a workflow re-saved with new step rows keeps its takes together), or step:{stepId:N} when the workflow or step is gone (OutputsController.TakeGroupKey).

The latest take is the newest render of a run that has ENDED. A render of a run still queued or running is never the latest take — it may still be replaced — and is never deleted. A render that is not the last one its own run produced for that step is intermediate (a review-loop pass or a retry the run then replaced).

API (/api/v1/projects/{projectId}/outputs, owner-scoped like every project endpoint: unknown project 404, another owner's 403):

MethodPathWhat it does
GET/outputsEvery render, newest first, now also carrying takeGroupKey, takeIndex (1 = latest), takeCount, isLatestTake, isIntermediate, isPinned — appended to OutputVideoResponse, every older field unchanged
POST/outputs/{stepResultId}/pin{ "pinned": true } pins a take (WorkflowStepResult.IsPinned, column is_pinned, engine migration AddStepResultIsPinned); a pinned take is kept by every cleanup
DELETE/outputs/{stepResultId}Deletes one earlier take; 409 for the latest take, a pinned take, or a take whose run is still going
POST/outputs/prune{ "takeGroupKey": "…", "keepLatest": 1 } ("Delete earlier takes") deletes every take of that video except the newest keepLatest (at least 1), the pinned ones and those of runs still going; returns { deleted, kept }

Deleting a take removes its stored object — only when no surviving take still uses the same key (a step-cache hit reuses an earlier run's object) — removes its project-file row if it was saved to Files, and clears the step result's OutputStorageKey. The step result itself stays, so the run's history (its output text, the EDL, the timing) is intact; the run page simply no longer shows a video for it.

Web. The project's Renders tab and the Overview's "Latest renders" show each video once, by its latest take (web/lib/outputs/takes.ts groupTakes, which falls back to workflow + step id against an older server). Under the player, the video being watched lists its earlier takes collapsed under "N earlier takes" (components/outputs/EarlierTakes.tsx): each playable and downloadable, marked "In-progress take" (intermediate) or "Run still going", pinnable, and deletable after a confirmation; "Delete earlier takes" asks once and calls prune. The Files tab hides the project-file rows of renders that are not the latest take of their video (earlierTakeStorageKeys — a key the latest take still plays from is never hidden) behind "Show N earlier takes of your videos". There is no automatic keep-the-last-N setting: prune with keepLatest is the building block for one.

See also One registered render per run, which stops the compile step from registering a Files entry per review-loop pass in the first place.


Scene/visual analysis (Phase 1)​

Phase 1 adds deterministic (no LLM, no new external dependency) visual and audio descriptors per shot on top of the shipped VideoAnalyze step, which until now only knew about cut points (silence gaps, shot-change timestamps) and speech. It is purely additive: nothing existing changes behavior when the new config defaults are used, other than the artifact and bounded view gaining new optional fields (VideoAnalysisArtifact.Version moves to 2, but a Version: 1 artifact still deserializes unchanged — every new field is optional/nullable/default-valued and appended, never inserted or reordered).

The technique: one low-res raw-frame grid pass​

A single new ffmpeg invocation decimates and downscales the source video to a raw RGB pixel grid file (FfmpegFrameGridSampler, FfmpegArgvBuilder.BuildGridSampleArgs):

ffmpeg -nostdin -hide_banner -y -loglevel error -protocol_whitelist file
-i {input}
-an -sn
-vf fps={sampleFps},scale={gw}:{gh}:flags=area,format=rgb24
-f rawvideo -pix_fmt rgb24
{scratch}/grid.rgb

flags=area is load-bearing — it's a true box-average downscale, so each output pixel is the exact mean of its source block, which is what makes per-region statistics meaningful. Default grid is 32x18 at 2.0 fps; MaxVisualSampleFrames (default 4000) clamps the effective fps downward for very long videos: effectiveFps = min(VisualSampleFps, MaxVisualSampleFrames / durationSec). Sample time of raw frame i is exactly i / effectiveFps (the fps filter emits CFR from t=0), so no timestamp parsing is needed anywhere downstream — just byte-offset math (frameSizeBytes = gridWidth * gridHeight * 3).

FrameGridAnalyzer (WorkflowEngine/Services/Video/FrameGridAnalyzer.cs) is a pure, unit-testable static class that derives everything below from this one grid buffer — no ffmpeg stderr scraping for any of it:

  • Motion — mean absolute luma delta between consecutive sampled frames within a shot, normalized 0..1: MotionMean/MotionPeak/MotionStdDev, bucketed into a MotionClass (Static/Subtle/Moderate/Dynamic).
  • Camera move — a 1-D SAD (sum-of-absolute-differences) integer pixel-shift search (dx/dy in [-4, 4]) between consecutive frames' column-sum and row-sum luma profiles. A consistent same-sign shift ⇒ Pan/Tilt; a high-variance alternating-sign shift ⇒ Handheld; more motion near the frame center than the border ⇒ Zoom; near-zero shift and near-zero motion ⇒ Static; otherwise Unknown. Reported as CameraMove + CameraMoveConfidence (0..1) — this is a documented heuristic, not ground truth, and is always paired with its confidence.
  • Still windows / head-tail motion — runs where per-frame motion stays below StillMotionThreshold for at least MinStillWindowMs (capped to MaxStillWindowsPerShot, longest kept), plus HeadMotion/TailMotion (mean motion over the shot's first/last 250ms) — "will a cut here land mid-motion?"
  • Exposure/color — Rec.709 luma per pixel, shot-averaged into BrightnessMean/ BrightnessStdDev (temporal flicker), ContrastRms (intra-frame luma std-dev, shot-averaged), ClippedHighlightRatio/CrushedBlackRatio, SaturationMean (HSV). DominantColors: the top 3 bins of a 64-bin (4 levels/channel) RGB histogram, each as a hex color + population share.
  • Regions / safe zones — the grid's 3x3 spatial cells (R0..R8) plus three named overlay-candidate bands (LowerThird, UpperThird, CenterBand). Per region: LumaMean, LumaStdDev (clutter proxy), TemporalMotion, TextColor (Light/Dark, from LumaMean), and Suitability (0..1) — 0.5*clutterScore + 0.3*motionScore + 0.2*extremeScore, favoring an uncluttered (low luma std-dev), low-motion region whose brightness isn't at a 0/1 extreme. BestOverlayRegion on the shot is the highest-Suitability name among the three named bands.
  • Near-duplicate / best-take grouping — per shot, a signature (FrameGridAnalyzer.ShotSignature): a z-normalized, time-averaged luma grid (LumaSig, 576 floats for 32x18) plus the L1-normalized 64-bin color histogram (ColorHist). Distance = 0.7*(1-cosine(LumaSig)) + 0.3*(0.5*L1(ColorHist)), Similarity = 1 - Distance. Groups form via single-linkage clustering over a sliding time window (DuplicateWindowShots, default 20 — only the previous N shots are compared, both faster and more correct since multi-take shots are temporally adjacent) at DuplicateSimilarityThreshold (default 0.90). Each group gets an id d{n}; members are ranked by a heuristic TakeQuality (documented in FrameGridAnalyzer.ComputeTakeQuality: 35% inverse motion jitter, 25% audio level, 20% exposure, 10% duration, 10% neutral sharpness placeholder) — rank 0 is IsBestTake. A singleton shot gets no group (DuplicateGroupId = null).
  • Ken-Burns candidate — KenBurnsCandidate = true when the shot is Static, MotionMean < 0.02, duration >= 2.5s, and at least one named band has decent Suitability; identifies candidates only — no zoompan is ever applied (out of scope, as always).
  • Audio levels — WavRmsSampler (pure, WorkflowEngine/Services/Video/WavRmsSampler.cs) windows the canonical 16kHz mono s16 WAV FfmpegAudioExtractor already produces into 250ms RMS/peak windows (20*log10(rms/32768), floored at -96 dBFS instead of -Infinity); reuses the WAV already extracted for transcription, or extracts it fresh if transcription is off/degraded. Per shot: AudioRmsDbfs (energy-weighted mean of overlapping windows), AudioPeakDbfs, and SpeechRatio (from the already-detected silence spans — no new audio pass), bucketed into a LoudnessClass (Quiet/Normal/Loud).
  • Not implemented in Phase 1 (opt-in, default false, lower priority than the core grid pipeline): DetectLetterbox/ActiveCrop and DetectSharpness/Sharpness — the config fields and artifact columns exist (always null/false) so a future phase can fill them in without another schema migration; MetadataPrintOutputParser/CropDetectOutputParser (generic ffmpeg ... metadata=print:file=- / cropdetect stderr parsers, mirroring SilenceDetectOutputParser/ShowinfoOutputParser's style) were likewise left unimplemented.

VisualDetail: degrade before drop​

VideoAnalyzeStepExecutor.BuildBoundedView treats "how many shots/silences/segments the model sees" as higher priority than "how much visual/audio detail each one carries". Before ever dropping an offered item to fit MaxOutputChars, it tries the configured VisualDetail level, then each lower level in turn — Full → Compact → None — re-serializing the same set of offered items at each level. Only once None (which renders a shot exactly as it looked before Phase 1 — no v/a key at all) still doesn't fit does the pre-existing item-dropping loop run. This guarantees meta.offeredIdCount can never be smaller than what a plain AnalyzeVisuals: false run would produce at the same MaxOutputChars — richer per-shot data can only ever cost detail, never cost coverage. meta.visual.detail (and the top-level meta.visualDetailApplied) records whichever level actually got used.

  • None — identical to the pre-Phase-1 shot shape.
  • Compact (default) — motion, move, cutIn/cutOut, the single longest still window, bright, contrast, the top 2 dominant colors, safe (best overlay region), dup/best, kenBurns (only when true), plus a: {rms, speech} when audio levels are available.
  • Full — everything Compact has, plus the full region list, the full still-window list, motionStdDev/motionPeak/cameraConfidence, and all (up to 3) dominant colors.

Every visual/audio number in the view is rounded before serialization — seconds to 2 decimal places, 0..1 scores to 0..100 integers (Round2/Score helpers) — since a raw double can serialize as 15-17 characters of floating-point noise, which adds up fast across hundreds of shots. The four base shot fields (id/startSec/endSec/durationSec) are deliberately left unrounded at every detail level, matching the pre-Phase-1 output exactly.

View/artifact shape additions​

{
"view": {
"shots": [{
"id": "s4", "startSec": 12.4, "endSec": 16.8, "durationSec": 4.4,
"v": {
"motion": 12, "move": "Pan", "cutIn": "still", "cutOut": "moving",
"still": [{ "startSec": 12.4, "endSec": 12.9 }],
"bright": 41, "contrast": 22, "colors": ["#2b3a4f", "#c9b48a"],
"safe": { "region": "LowerThird", "fit": 88, "text": "Light" },
"dup": "d2", "best": true, "kenBurns": true
},
"a": { "rms": -21, "speech": 82 }
}],
"pacing": { "meanShotSec": 4.1, "medianShotSec": 3.8, "cutsPerMinute": 14.6, "motionTimeline": [12, 30, 8], "timelineBinSec": 5.0 },
"duplicateGroups": [{ "id": "d2", "shotIds": ["s4", "s7"], "bestShotId": "s7", "similarity": 94 }]
},
"meta": {
"visual": { "applied": true, "degraded": false, "sampleFps": 2.0, "gridWidth": 32, "gridHeight": 18, "detail": "Compact" },
"audioLevels": { "applied": true }
}
}

Pacing (FrameGridAnalyzer.ComputePacing) is a whole-artifact summary — MeanShotSeconds/ MedianShotSeconds/CutsPerMinute from shot timing alone, plus a MotionTimeline (mean MotionMean per TimelineBinSeconds-wide bin, empty when no shot has visual data). pacing/ duplicateGroups are only surfaced in the view at Compact/Full detail — never at None — so that a budget-forced collapse to None (whether from AnalyzeVisuals: false or from a degraded visual-analysis stage) is always byte-identical regardless of whether visual data merely got suppressed by degradation vs. never computed at all.

Failure handling: degrade, never fail the step​

Both the visual-analysis block (grid sampling + every FrameGridAnalyzer call) and the audio-level block (WAV sampling) are wrapped in their own try/catch, exactly like the existing transcription-degrade pattern: on any exception, Provenance.VisualAnalysisApplied/ AudioLevelsApplied become false, VisualAnalysisDegraded becomes true, a warning is logged, and the step continues with shots/silences/transcript exactly as if that stage were configured off. Nothing in Phase 1 can fail a VideoAnalyze step.


Vision captioning (Phase 2)​

Optional vision-LLM shot captioning on top of Phase 1's deterministic descriptors: extract one representative keyframe per selected shot, send it to a vision-capable chat model, and get back a short structured scene description (subjects, action, setting, mood, shot scale, camera angle, on-screen text, tags). Unlike every other stage in this document, this one is off by default and makes real LLM calls — see Why Off, not Optional below.

InferenceProviderCapability.Vision​

A third InferenceProvider capability, alongside Chat and Transcription. It reuses the exact same chat-completions machinery Chat does — IChatClientFactory/ResolvedInferenceProvider gained no new members — since a vision call is just an ordinary chat-completions call with an image content part alongside the text prompt. It is still a separate capability (not folded into Chat) so a vision-capable deployment (which may differ from the deployment an agent's chat resolution uses) can be configured and defaulted independently: Vision participates in its own "at most one default" bucket via the same composite (capability, is_default) partial unique index Chat/Transcription already share — no schema migration was needed to add the third value, since that index is generic over any capability value, not hardcoded to two.

IInferenceProviderResolver.ResolveVisionAsync(explicitProviderId, ct) mirrors ResolveTranscriptionAsync's precedence and null-when-nothing-resolves contract exactly (explicitProviderId if it resolves to an enabled Vision-capability row → the single enabled IsDefault && Capability == Vision row → null), but returns ResolvedInferenceProvider? (not a dedicated vision type) and resolves through the same ResolveFromProvider helper ResolveAsync (chat) uses. Like transcription, there is deliberately no fallback to the legacy AzureOpenAI:* config keys — silently sending an image to a deployment that may not support vision would fail confusingly.

POST /api/v1/inference-providers/{id}/test and POST /api/v1/inference-providers/test gained a Vision arm alongside the existing Transcription arm: it sends a trivial embedded 1x1 JPEG through IChatClientFactory with a "reply with the word ok" prompt and treats any non-empty response as success — mirroring the existing synthesized-silent-WAV transcription ping. The admin UI (InferenceProviderForm) exposes Vision as a third Capability option; the per-agent provider override picker (AgentInferenceProviderSelect) filters to Chat rows only (an allowlist, not merely "not Transcription" — a denylist would have silently admitted Vision rows here too), since that override feeds chat resolution only.

Why Off, not Optional​

VideoAnalyzeStepConfig.Transcription defaults to Optional because it costs at most one ASR network call per step. Captioning is structurally different: it can cost up to MaxCaptionedShots (default 24) separate vision chat-completion calls — each carrying an image — per VideoAnalyze step. If Vision defaulted to Optional, then the moment any admin configured a Vision-capability default provider for some unrelated workflow that actually wants captioning, every other existing or future VideoAnalyze step in the system would silently start making real, billed vision calls with no config change of its own. Defaulting to Off keeps every step's cost/latency unchanged unless its author explicitly opts in by setting Vision on that step. Optional/Required otherwise carry the same degrade-vs-fail semantics Transcription does.

Keyframe selection​

KeyframeSelector (WorkflowEngine/Services/Video/KeyframeSelector.cs) is pure — no ffmpeg, no I/O:

  • ChooseKeyframeSec(shot) — the midpoint of the shot's longest StillWindow (Phase 1) when at least one exists (a calm moment makes a cleaner, less motion-blurred frame), else the shot's own midpoint. Falls back to the shot midpoint when Visual is null (visual analysis off, degraded, or simply no still windows).
  • SelectShotsToCaption(shots, duplicateGroups, strategy, maxCaptionedShots, minCaptionShotSeconds) implements three strategies (VideoCaptionSelection), excluding any shot shorter than minCaptionShotSeconds, and always returning ids in shot-chronological order regardless of selection order:
    • PerDuplicateGroup (default) — captions each near-duplicate group's best-take shot first (Phase 1's DuplicateGroups), so N takes of one setup cost one vision call, not N; fills any remaining budget with the longest not-yet-selected shots.
    • LongestShots — simply the N longest eligible shots.
    • EvenlySpaced — shots at roughly even index intervals across the whole shot list.

Keyframe extraction​

FfmpegArgvBuilder.BuildKeyframeArgs(inputPath, outputJpgPath, atSec, maxWidth) — a new ffmpeg invocation alongside Phase 1's grid-sample/audio-extract builders, following the exact same IVideoToolRunner calling convention via a new IKeyframeExtractor/FfmpegKeyframeExtractor pair (mirroring IAudioExtractor/FfmpegAudioExtractor):

ffmpeg -nostdin -hide_banner -y -loglevel error -protocol_whitelist file
-ss {atSec} -i {inputPath}
-frames:v 1
-vf scale='min({maxWidth},iw)':-2
-f image2 -c:v mjpeg -q:v 4
{outputJpgPath}

-ss before -i for fast input seeking (same convention as BuildExtractAudioArgs); the scale expression never upscales and preserves aspect ratio. IKeyframeExtractor is deliberately its own small interface (not folded into IFrameGridSampler) so Phase 3 (motion-graphics overlay planning, applied at compile time) has a single established pattern to follow for its own new ffmpeg operations.

The captioner​

IShotCaptioner/VisionShotCaptioner (WorkflowEngine/Services/Video/) build one chat message per shot — a text prompt plus the keyframe JPEG as a DataContent("image/jpeg") content part — and call IChatClient.GetResponseAsync with ChatResponseFormat.ForJsonSchema<VideoShotCaption>(), the exact same structured-output mechanism AgentStepExecutor/ReelBoltAgentBase use for agent steps. IChatClientFactory is reused unchanged.

public sealed record VideoShotCaption(
string ShotId, string Summary, IReadOnlyList<string> Subjects,
string Action, string Setting, string Mood,
string ShotScale, string CameraAngle,
IReadOnlyList<string> OnScreenText, IReadOnlyList<string> Tags);

Critical safety property: the shot-id ↔ caption binding is never model-controlled. The model is called once per shot; VideoAnalyzeStepExecutor — never the captioner — always overwrites the returned VideoShotCaption.ShotId with the id it actually requested (ShotCaptionRequest.ShotId) before attaching the caption to a shot. This mirrors the id-anchored discipline VideoEditDecisionOutput/VideoCompileStepExecutor already use: never trust an identifier the model echoes back for anything that matters. Covered by a dedicated test asserting the executor ignores a deliberately-wrong model-returned ShotId.

Captioning retries up to 2 attempts per shot with a short linear backoff (mirrors TranscribeWithRetryAsync's shape); a captioning failure on one shot never aborts captioning of the rest — the executor's per-shot loop catches, counts it in meta.vision.failedShots, and moves on.

Executor wiring​

Captioning runs last among VideoAnalyze's analysis stages, strictly after every deterministic stage (silence/shot detection, transcription, Phase 1 visual/audio analysis, near-duplicate grouping) — so a vision failure can never put anything deterministic at risk:

Multi-source hoist (Phase 4 fix). Captioning is a step-level pass over every source's shots combined, run once after all per-source deterministic analysis completes — not a per-source pass run once per source. This matters for MaxCaptionedShots/VisionTimeoutSeconds: both are genuinely step-wide budgets across every source clip, never silently reset per source. An earlier draft of this feature ran captioning inside the per-source loop, which would have let a 2-source analysis spend up to 2 × MaxCaptionedShots vision calls instead of the configured cap — caught and fixed before merge; VideoMultiSourceTests pins the corrected step-wide behavior.

  1. Vision == Off → skipped entirely, meta.vision = {mode: "Off", applied: false, ...}. Every other field in the view/artifact is byte-identical to a pre-Phase-2 run — the single most important regression test in this phase, mirroring Phase 1's degrade-before-drop test's importance.
  2. Else, resolve via ResolveVisionAsync. null + Required ⇒ fail with VISION_UNAVAILABLE. null + Optional ⇒ degrade (meta.vision.degraded = true), continue with no captions.
  3. Resolved ⇒ select shots via KeyframeSelector, extract each keyframe to scratch, caption each with the 2-attempt retry. An aggregate VisionTimeoutSeconds wall-clock budget covers the whole captioning pass (not per-shot) via a linked, timed CancellationTokenSource; exceeding it mid-pass stops captioning further shots, keeps whatever already succeeded, and sets meta.vision.partial = true.
  4. If zero captions were obtained after all that: Required ⇒ fail with VISION_FAILED; Optional ⇒ degrade and continue with zero (or partial) captions.

View/artifact shape addition​

A shot with a caption gains a "c" key in the bounded view, gated on VisualDetail exactly like Phase 1's "v"/"a" keys — same degrade-before-drop discipline, never bypassed:

{
"view": {
"shots": [{
"id": "s4", "startSec": 12.4, "endSec": 16.8, "durationSec": 4.4,
"c": {
"summary": "A presenter gestures at a whiteboard while explaining a diagram.",
"scale": "Medium", "mood": "Focused", "tags": ["presenter", "whiteboard", "explaining"],
"subjects": ["presenter"], "action": "gesturing at a diagram",
"setting": "office whiteboard", "cameraAngle": "Eye level", "onScreenText": [],
"style": "flat ungraded log", "issues": ["soft focus"]
}
}]
},
"meta": {
"vision": { "mode": "Optional", "applied": true, "provider": "gpt-4o-mini-vision", "degraded": false, "partial": false, "captionedShots": 6, "failedShots": 0, "persistedKeyframes": 0 }
}
}

Compact detail shows summary/scale/mood/tags/style/issues (Phase 4 added style and issues at Compact, matching Phase 1's own "cheap signals first" discipline); Full adds subjects/action/setting/cameraAngle/onScreenText/timeOfDay/lighting/framing (Phase 4). A shot with no "c" key is normal — not selected for captioning, or captioning off/failed/ degraded — never a signal the shot is empty or unimportant; the VideoStoryEditor prompt says so explicitly.

VideoAnalysisArtifact.Version stays at 2, not 3: Phase 2 appends exactly one more optional/nullable field (VideoAnalysisShot.Caption) plus optional/default-valued fields on VideoAnalysisProvenance, and a Version-2-without-captions artifact and a Version-2-with-captions artifact are both valid under the identical shape — no consumer needs to structurally distinguish them (a consumer that cares simply checks whether Caption is null). Phase 4's five new VideoShotCaption fields (TimeOfDay, Lighting, VisualStyle, Framing, TechnicalIssues) are additive the same way and do not bump the version either.

Prompt priming, contact sheets, and persisted keyframes (Phase 4)​

Three changes to the captioning path itself, none of which touch the id-anchored/no-timestamp contract:

  • Prompt priming. The vision prompt is primed with a short sentence of Phase 1's own deterministic measurements for the shot being captioned (color temperature, tone curve, saturation, camera move, letterbox/backlit flags) — words derived from measurements, never a number the model could restate. The model is told to use this only to inform its reading, not to treat it as ground truth it must repeat; null (visual analysis off/degraded for that shot) omits the block entirely. ShotCaptionRequest.MeasuredContext carries this; VisionShotCaptioner formats and injects it.
  • Contact-sheet keyframes (KeyframesPerShot, default 1). 1 extracts the same single mid-shot still as before Phase 4 — byte-identical. 2 or 3 instead extracts that many frames evenly spaced across the shot (KeyframeSelector.ChooseKeyframeSecs) and hstacks them into one contact-sheet JPEG (FfmpegArgvBuilder.BuildContactSheetArgs, IKeyframeExtractor.ExtractContactSheetAsync) — N cheap input seeks, never a single pass decoding the whole shot. Each pane is scaled to KeyframeMaxWidth / N so the combined image and its vision-call token cost stay roughly flat. The prompt gains one extra sentence telling the model the image is a multi-frame contact sheet of one shot, not several shots.
  • PersistKeyframes is now wired: true uploads each captioned shot's keyframe JPEG (or contact sheet) to storage under video-analysis/{executionId}/step-{stepOrder}-keyframes/{shotId}.jpg, counted in meta.vision.persistedKeyframes. A persist failure is logged and swallowed — it can never cost the caption itself, since observability is never allowed to be more load-bearing than the thing it observes. Default stays false (scratch-only, deleted with the rest of scratch space).

Also (judgment call 9): KeyframeSelector's longest-shots fill pass now round-robins its selection across source clips (grouping candidates by SourceIndex, longest-first within each group, then alternating groups) instead of picking the global longest shots regardless of source — so a MaxCaptionedShots budget spent across several source clips isn't silently exhausted entirely on one long clip. A single-source analysis produces exactly one round-robin "group", which is provably identical to the pre-Phase-4 plain longest-first order — no behavior change for the (still overwhelmingly common) single-source case.

Builder UI for Phase 2 and Phase 4​

The per-step VideoAnalyze config panel now has a "Vision captioning" section covering Vision, VisionProviderId (filtered to Vision-capability provider rows by the same allowlist discipline the per-agent chat override uses — never "not Transcription"), CaptionSelection, MaxCaptionedShots, KeyframesPerShot, KeyframeMaxWidth and PersistKeyframes, alongside the three free Phase 4 switches — see Semantic visual dimensions (Phase 4). The remaining fine-tuning fields (MinCaptionShotSeconds, VisionTimeoutSeconds, MaxCaptionChars, MaxSharpnessShots, the look-grouping thresholds) stay settable via the raw step JSON only — they are cost/quality trim, not pipeline shape.


Transcription (ASR)​

Ships in full — not deferred — via the same InferenceProvider model chat completions already use, distinguished by a new Capability column (Chat or Transcription). A single provider row can't serve both roles: a Whisper deployment is a different deployment from a chat deployment, and many OpenAI-compatible chat gateways have no /audio/transcriptions endpoint at all (only whisper.cpp-server/faster-whisper-server/speaches/LiteLLM-style deployments do). Chat and Transcription each have their own independent "at most one default" constraint (a composite unique index on (capability, is_default)), and every resolution path — chat and transcription — filters on its own Capability explicitly.

VideoAnalyzeStepConfig.Transcription has three modes:

  • Off — silence + scene + loudness only. Deterministic, zero external dependency, always available.
  • Optional (default) — attempt ASR; if no Transcription-capability provider resolves, or the call fails after a small bounded retry, degrade cleanly to Off and record meta.transcription.degraded = true in the bounded view.
  • Required — fail the step with a precise diagnostic if ASR is unavailable, rather than silently shipping an edit decision with no transcript context.

Long audio is chunked at silence-boundary-aligned points (never mid-word) by the pure, independently-tested TranscriptChunkPlanner, and each chunk's word/segment timestamps are offset by that chunk's absolute start before concatenation — the single most likely correctness bug in this feature class, and the reason it has a dedicated multi-chunk offset test.

The admin UI for configuring providers (/admin/inference-providers) exposes Capability as a field on create/edit; the per-agent chat-provider override picker filters Transcription rows out, since that override only ever feeds chat resolution.


Transcripts from clips with no speech​

Whisper-family recognisers fed music, room tone or silence routinely "hear" stock phrases — "Thank you for watching.", "you", a subtitle credit. Before this gate a VideoAnalyze step offered such a line to the story editor as real dialogue (a customer's b-roll edit chose its closing shot because of one). TranscriptSpeechGate (WorkflowEngine/Services/Video/TranscriptSpeechGate.cs, pure) now discards those segments per source clip, before t{n}/w{n} ids are assigned, so the id namespaces stay contiguous and a dropped segment can never be offered, quoted to an agent, screened, or resolved by VideoCompile. A segment is dropped when any of these holds:

  • no_speech — the recogniser says so itself: Whisper's own rule, no_speech_prob ≥ 0.6 with avg_logprob < -1 (or no log-prob reported), or no_speech_prob ≥ 0.9 outright. The figures come from verbose_json (TranscriptSegment.NoSpeechProb/AvgLogprob/CompressionRatio, optional members read by both transcription clients; a backend that omits them simply skips this rule).
  • repetition — compression_ratio > 2.4, Whisper's looping-output signature.
  • no_speech_energy — the step's own measurements: silence detection covers ≥ 80% of the segment's window, or no Phase 1 audio-level window under it rises above SilenceThresholdDb.
  • stock_phrase — the whole segment is one of a short list of phrases recognisers invent on non-speech audio, and nothing vouches for it: the recogniser reported no_speech_prob ≥ 0.2, or reported no confidence at all and the segment is the clip's only one. A confidently recognised "thank you" inside real dialogue is kept.

Words whose midpoint falls inside a dropped segment (and no kept one) go with it. A clip with no audio stream never reaches the gate: it is not transcribed at all, so it gets no transcript.

What was dropped is recorded, never silent: meta.transcription.droppedNoSpeech (present only when non-zero, so a clean transcript's meta is unchanged) and VideoAnalysisProvenance.TranscriptSegmentsDroppedNoSpeech in the artifact, summed across sources. The VideoAnalyze revision tag in DeterministicStepRevisions moved with it, so cached analyses are recomputed. Tests: TranscriptSpeechGateTests (each rule) and the "Transcript speech gate" block of VideoAnalyzeStepExecutorTests (the hallucinated outro never reaches the view; t ids stay contiguous; a no-audio clip is never sent to the recogniser).

Motion graphics (Phase 3)​

Optional motion-graphics overlays — lower-thirds, titles, callouts — applied during VideoCompile's encode, planned by a third built-in agent that reasons over overlay-safe-zone "placement" candidates derived deterministically from Phase 1's per-shot region data, but never emits a timestamp or pixel coordinate — the same structural discipline VideoEditDecisionOutput/VideoStoryEditorAgent already established, extended to cover geometry as well as time. Off by default (VideoAnalyzeStepConfig.EmitOverlayPlacements = false, VideoCompileStepConfig.EnableGraphics = false) — both flags must be explicitly opted into, and EnableGraphics = false leaves the compile path byte-identical to the pre-Phase-3 behavior.

This is the highest-risk phase of the feature: motion-graphics overlay TEXT is the first model-authored content in this feature to reach ffmpeg at all (every prior stage passes only opaque ids and workflow-author-supplied enum/allowlisted config). The text-sanitization and textfile-based discipline below is the core deliverable of this phase, not polish on top of it.

Placement candidates (deterministic, in VideoAnalyze)​

OverlayPlacementBuilder (WorkflowEngine/Services/Video/OverlayPlacementBuilder.cs) is a pure, unit-tested static class deriving overlay-placement candidates from shots that already carry Phase 1 Visual/Regions data (nothing is computed if visual analysis is off or degraded). For each shot: rank its three NAMED overlay-safe-zone bands — LowerThird/UpperThird/CenterBand (never the 3x3 R0..R8 grid cells, which exist for other diagnostics, not placement) — by Suitability descending, take up to MaxPlacementsPerShot, and resolve a time window:

  1. LongestStillWindow — the shot's longest Phase 1 still window, if one exists (clamped to the shot's own bounds).
  2. ShotMiddle — else, a ~2.5s window centered on the shot's midpoint, clamped to the shot's own bounds.
  3. ShotStart — else (a degenerate ShotMiddle window, e.g. an extremely short shot), a window anchored at the shot's start, clamped to the shot's own bounds.

The whole artifact is capped at MaxPlacements, dropping the lowest-suitability candidates first; survivors are then re-sorted into deterministic generation order (shot order, then per-shot rank) and assigned SEQUENTIAL, globally unique ids p0, p1, … — a single counter across the whole artifact, not per-shot.

Critical id-isolation requirement: p{n} placement ids are a SEPARATE namespace from s{n}/g{n}/t{n} cut-anchor ids. VideoAnalysisArtifact.OfferedPlacementIds is a SEPARATE list from OfferedIds, gated on the exact same "degrade before drop" VisualDetail discipline Phase 1 established: view.placements is shown as one atomic array at Compact/Full detail and omitted entirely at None — so a placement id is only ever "offered" when the whole array survived to whatever detail level the view actually settled on. VideoCompileStepExecutor.BuildIdTimeIndex (which resolves Keep span ids to cut times) deliberately does NOT include placements — a code comment there explains why — and a dedicated test proves a Keep span naming a p0 id fails UNKNOWN_ID exactly like any other id that index does not contain.

view.placements entries are deliberately small: {id, shotId, region, startSec, endSec, fit, text} (plus src in a multi-source run) — a 2dp-rounded source-timeline window, a 0-100 suitability score, and a "Light"/"Dark" text-color hint, never a rect (geometry is resolved server-side only, at compile time, from the full artifact).

When the MotionGraphicsPlanner step runs with agentInputContextMode: FullWorkflow (the only mode that gives it the story editor's decision at all), each placement also gets an inEdit: bool — whether that candidate's own moment survives the cut, computed by MotionGraphicsPlacementAnnotator at prompt-assembly time in AgentStepExecutor (display-only: nothing is filtered, an inEdit: false candidate stays fully choosable) by resolving the kept spans through the exact same VideoCompileStepExecutor.MapSourceWindowToOutput that later drops an overlay with reason cut_away — so the planner sees the survival verdict BEFORE it picks, instead of discovering it only after the fact at compile time. Absent entirely under PreviousStepOnly/other context modes, since there's no decision yet to check against.

Sentence-boundary and punctuation-reliability signals. Every view.segments transcript segment also carries endsSentence: bool — whether that segment's own text (its full, untruncated form, not the MaxSegmentTextChars-shortened copy the view shows) ends in terminal sentence punctuation. This is the authoritative signal AgentType.VideoStoryEditor is told to use instead of eyeballing the text field itself for a mid-sentence cut. But punctuation output is only as trustworthy as the ASR backend that produced it — real deployments have been observed transcribing some clips with almost no terminal punctuation at all even though every answer is a complete thought — so meta.transcription.punctuated carries one entry per source clip, {src, segments, punctuatedSegments, ratio, reliable} (reliable is ratio >= 0.5), and the prompt is told to fall back to judging sentence-completeness semantically from the segment text whenever the relevant source's entry says reliable: false. The same ratio/reliable pair (recomputed for the specific source the compile step's sentenceCheck is about, not the whole run) is surfaced to AgentType.VideoReviewAgent too, whose otherwise-hard "ends mid-sentence" score cap is downgraded to a soft, non-blocking observation when that source's signal is unreliable — a punctuation-starved ASR result must not by itself force a low score on an otherwise-good edit.

startSec/endSec are input to the planner, not output from it, and that distinction is the whole no-timestamp rule: MotionGraphicsPlanOutput still has no time-bearing property whatsoever (MotionGraphicsPlanOutputInvariantTests), and VideoCompileStepExecutor still resolves an overlay's real window from the full artifact's own unrounded VideoAnalysisPlacement.StartSec/EndSec — never from the rounded numbers the view showed. The precedent is view.segments, which has exposed exactly these two field names to AgentType.VideoStoryEditor since the feature's first phase.

They are there because without them the view was undecidable. MaxTimeSlicesPerRegion splits one long shot's region into several candidate sub-windows, so on a static shot (a talking-head interview, say) several placements serialized byte-identically except for their id:

{"id":"p6","shotId":"s1","region":"LowerThird","fit":69,"text":"Light"}
{"id":"p7","shotId":"s1","region":"LowerThird","fit":69,"text":"Light"}

The planner had no information on which to prefer one over another and defaulted to the first of each identical run — an arbitrary choice dressed up as an editorial one. The time window is the only thing that distinguishes them, and it is also what lets the planner line an overlay up with the view.segments entry (same units, same source timeline) whose content the overlay is about.

AgentType.MotionGraphicsPlanner and MotionGraphicsPlanOutput​

An ordinary StepType.Agent step — no new step type, exactly the precedent AgentType.VideoStoryEditor set. Given the story editor's decision (or the same bounded view) plus view.placements, it decides zero or more overlays:

public class MotionGraphicsOverlay
{
public string PlacementId { get; set; } = ""; // must be in OfferedPlacementIds
public string Kind { get; set; } = ""; // LowerThird | Title | Callout | Tag
public string Text { get; set; } = "";
public string Subtext { get; set; } = "";
public string Duration { get; set; } = ""; // Short | Medium | Hold — never a number
public string Emphasis { get; set; } = ""; // Subtle | Normal | Strong
public string Reason { get; set; } = "";
}

public class MotionGraphicsPlanOutput
{
public List<MotionGraphicsOverlay> Overlays { get; set; } = new();
public string PlanRationale { get; set; } = "";
}

Every property on both types is a string/List<string-bearing-type> — the SAME rushcut invariant VideoEditDecisionOutput established, extended to also forbid a pixel coordinate. MotionGraphicsPlanOutputInvariantTests mirrors VideoEditDecisionOutputInvariantTests's reflection approach exactly. Duration/Emphasis/Kind are enum WORDS the model chooses from a closed vocabulary described in its prompt — VideoCompileStepExecutor alone resolves Duration to milliseconds (OverlayShortMs/OverlayMediumMs/OverlayHoldMs, defaulting to Medium on an unrecognized value) and OverlayFontSizePct to an actual pixel font size, which Emphasis (Subtle/Normal/Strong) then nudges up or down by a fixed multiplier (0.8x/1x/1.25x, still clamped to the same valid [2, 12] percentage range) in DrawtextFilterBuilder.ComputeFontSize — a deliberately small effect, not a whole per-emphasis styling system. Kind (LowerThird/Title/Callout/Tag) is currently descriptive/reserved only — nothing reads it downstream, every kind renders identically, since the placement's own region (LowerThird/UpperThird/CenterBand) already resolves the overlay's geometry and letting Kind also influence position would create two disagreeing sources of geometry for the same overlay. See the doc comment on MotionGraphicsOverlay.Kind in OutputSchemas.cs. The MotionGraphicsPlannerAgent class's fallback prompt and the seeded built-in AgentDefinition row in DatabaseSeeder are kept verbatim-identical, enforced by the second [Fact] in VideoStoryEditorPromptConsistencyTests.cs (MotionGraphicsPlanner_fallback_prompt_matches_the_seeded_built_in_agent_prompt_verbatim, mirroring the first fact's reflection approach for VideoStoryEditorAgent). Tool access is now the same full sandbox+Remotion+render pipeline AuthorAgent gets, minus WriteProjectFile (AgentToolProvider) — widened from the minimal read-only scope VideoStoryEditor gets, since this agent can optionally back an overlay with a real, rendered Remotion asset (see RenderedAssetStorageKey below) rather than only plain drawtext/drawbox text. This is the only other agent besides AuthorAgent granted RenderVideoAndUploadToStorage, and it is a conscious tradeoff: this agent's prompt includes analysis-view content derived from the source video itself (on-screen text the Phase 2 vision model read, ASR transcript text), so it is the first agent in this feature with code-execution tools whose prompt is not limited to user-selected project files. The sandbox's own containment (read-only rootfs, no network egress by default, no Docker-socket access — see Security) is what bounds the blast radius of a successful prompt injection here: worst case is sandbox-contained code execution, not host compromise. See CLAUDE.md's "Agent Types (enum)" section for the same note.

Rendered-asset overlays​

An overlay can optionally carry MotionGraphicsOverlay.RenderedAssetStorageKey — the S3 storage key of a transparent-background motion-graphics asset the MotionGraphicsPlanner agent produced ITSELF, by actually calling RenderVideoAndUploadToStorage (a real tool call performing a real Remotion render and a real upload), rather than an id merely echoed back from a set the model was shown. Still not trusted blindly: VideoCompileStepExecutor validates the key matches the exact projects/{projectId}/outputFiles/{executionId}/... prefix RenderVideoAndUploadToStorage itself constructs for the CURRENT execution, before downloading or compositing anything at that key.

A rendered asset is the DEFAULT, not an upgrade. The tool grant that gives MotionGraphicsPlanner/MotionGraphicsDirector the full sandbox+Remotion+render set exists for this and nothing else, and the fallback a plain-text overlay reaches is drawtext over drawbox — type on a translucent rectangle, visibly cheaper than anything the agent could design. Both prompts originally framed rendering as "preferred when the moment deserves it" and plain text as "the simple fallback"; run for real, the graphics room rendered one asset out of three overlays on one pass and none out of three on the next, with no failed attempt in either — it simply took the cheaper branch. The prompts, the room charter and the CopyArtist seat persona now all state the same thing: every overlay is a designed render unless that specific moment has a specific reason for unadorned type, stated in the overlay's own reason. A render that fails after its retry budget may still fall back to text for that ONE overlay.

When present, OverlayAssetFilterBuilder (WorkflowEngine/Services/Video/OverlayAssetFilterBuilder.cs) composites the asset via ffmpeg's overlay filter — added as an extra input, stretch-scaled to the same compact ACCENT box geometry DrawtextFilterBuilder computes for a text overlay at the same placement (DrawtextFilterBuilder.ComputeAccentBoxPixels, shared by both builders), and time-shifted with setpts so the asset's own frame 0 lands at the overlay's actual on-screen start time on the OUTPUT timeline. eof_action=pass means once the asset's content runs out the overlay simply stops contributing (reverts to the plain cut) rather than freezing on its last frame for a longer "Hold" window. Text/Subtext are ignored for an overlay that carries a rendered asset — an overlay is one or the other, never both; a workflow author combines a rendered graphic with separate caption text by authoring two overlays at different placements. Empty/absent (the default) leaves an overlay a plain text overlay exactly as before this field existed — additive, not a replacement.

Alpha channel is validated, never trusted. Two independent guards close the "a real, transparent Remotion render still comes out as a solid, opaque block over the edited video" failure mode (observed in production as a black square with giant drop-shadow silhouettes stamped into it):

  1. Server-side validation — after downloading and ffprobing the asset, VideoCompileStepExecutor checks the probed pix_fmt against AlphaPixelFormats.HasAlpha (an allowlist of alpha-carrying formats — yuva420p/yuva444p10le/rgba/argb/gbrap/... — same allowlist discipline as the codec/color allowlists elsewhere in this file). An asset with no real alpha plane — e.g. the agent ignored the documented --pixel-format=yuva420p --codec=vp9 render recipe and produced a plain H.264 file — is dropped for THIS overlay only (droppedOverlays reason asset_missing_alpha_channel), the same per-overlay degrade-not-fail discipline as a corrupt or unprobeable asset (asset_download_or_probe_failed), never composited opaque.
  2. Filtergraph format-pinning — even when the source genuinely carries alpha, OverlayAssetFilterBuilder prepends format=rgba as the FIRST operation on the overlay's own input, before scale. This guards against a well-known libavfilter gotcha: format negotiation between scale and the downstream overlay filter can silently agree on a non-alpha common pixel format even when the decoded input has a real alpha plane, discarding transparency with no error of any kind. Pinning the format immediately after decode, before any negotiation happens, is what makes a genuinely transparent render actually composite transparently.

The source-to-output timeline mapping problem​

A placement's window was resolved against the SOURCE video during VideoAnalyze. But drawtext's enable=/alpha= expressions run against the ffmpeg filtergraph's OUTPUT timeline — the one the existing select/setpts cut stage produces, which is shorter than the source and has every cut gap removed entirely. Two internal, directly-unit-tested static functions on VideoCompileStepExecutor solve this purely from the already-resolved, frame-quantized ResolvedSpan list (their SnappedStart/SnappedEnd — the ACTUAL output-determining times, never the pre-quantization requested ones):

internal static double? MapSourceToOutputSec(IReadOnlyList<ResolvedSpan> spans, double sourceSec);

internal static (double Start, double End)? MapSourceWindowToOutput(
IReadOnlyList<ResolvedSpan> spans, double startSec, double endSec);

Both walk spans accumulating output-timeline duration as they go. MapSourceToOutputSec returns null when the second falls inside a cut gap (no corresponding output frame exists). MapSourceWindowToOutput intersects a window with the kept spans and returns the output-timeline window for the FIRST kept portion it overlaps (a single on-screen overlay cannot span a gap in the output video) — null if the window never overlaps any kept span at all. Resolving a placement's on-screen window intersects FIRST, truncates SECOND: the full candidate window [placement.StartSec, shot.EndSec] (clamped to the OWNING SHOT's own bounds — looked up via placement.ShotId against the full artifact's Shots, not merely the placement's own already-narrow window) is mapped through MapSourceWindowToOutput to find where it survives the cut, and only THEN is the model's chosen Duration applied, from wherever that surviving intersection begins. Truncating to durationMs before intersecting (the original, buggy order) silently dropped overlays whose full window overlapped a kept span by many seconds but whose first durationMs alone did not — verified live: a placement with real window [0, 29.83] and a kept span [4.8, 51.4] was tested only against [0, 3.0] (entirely cut) and wrongly dropped as cut_away.

Text sanitization and the textfile=/expansion=none discipline​

This is what makes the existing "not one model-originated character reaches an ffmpeg argv" claim (see Security) need a qualification, not a retraction. Overlay text is model-authored and does reach ffmpeg for the first time in this feature — but only as sanitized FILE CONTENT, never as argv or filter-string content:

  1. OverlayTextSanitizer.Sanitize(raw, maxChars) (WorkflowEngine/Services/Video/) — an ALLOWLIST (never a denylist, which is only ever safe against characters someone thought of) of letters, digits, combining marks, spaces, and a small safe punctuation set. NFC-normalizes first; collapses ALL whitespace (including newlines/tabs — drawtext treats a raw newline as a forced line break) to single spaces before the allowlist strips anything, so a newline becomes a space rather than being silently deleted (which would wrongly glue two words together); truncates to maxChars without splitting a grapheme cluster (StringInfo-based, never a blind str[..n]); returns "" for an empty/whitespace-only result, and the caller then drops that overlay/line entirely. The allowlist is a second, independent layer, not the primary safety mechanism — the primary mechanism is architectural (point 2 below): sanitized text is never interpolated into a filter/argv string at all, so no character reaching this far could ever terminate a drawtext option or invoke an expansion regardless of what the allowlist admits. Given that, the allowlist is kept narrow anyway, as ordinary defense-in-depth: colon (:) and percent (%) are excluded (drawtext's own option separator / expansion syntax), and so are ', ,, and ; (not filter-syntax-significant inside a quoted value, but not essential to a lower-third/title/callout either, so excluding them keeps the surface small). Emoji are also deliberately excluded (outside \p{L}/\p{N}/\p{M}): the configured overlay font is not guaranteed to carry emoji glyphs, so admitting them risks silent tofu-box rendering. Combining marks (\p{M}) ARE admitted so NFC-normalized text in scripts without precomposed forms (e.g. Devanagari vowel signs) survives sanitization instead of being silently mangled character-by-character.
  2. Text never appears in the ffmpeg argv or filter string at all. Each overlay's sanitized text/subtext is written to its own scratch file ({scratch}/ov-{slot}.txt — DrawtextFilterBuilder.MainTextSlot/SubtextSlot are the single source of truth both the writer and the filter-string builder use for slot numbering) and referenced via drawtext's textfile= option, never an inline text= value — so no drawtext metacharacter (:, ', \, %) in the text can ever terminate or inject into the filter string, because the text literally never appears in that string.
  3. expansion=none on every drawtext filter — disables drawtext's own %{...} expansion syntax (which can read pts/localtime/metadata or run %{eif:...} expressions) as defense-in-depth, even though the text is already sanitized and file-based.
  4. Every geometry/timing number is computed in C# (DrawtextFilterBuilder, from the probed frame size and the already-resolved output-timeline window) and formatted via FfmpegArgvFormat.Number — culture-invariant, exactly like EncodeReencodeAsync's existing BetweenTerms() — before being interpolated into the filter string. The model never supplies any of this directly; its only contributions are an offered placement id and a handful of already-resolved enum words.

DrawtextFilterBuilder.BuildFilterChain renames the cut stage's output label from [vout] to [vcut] only when there is at least one overlay to draw (when EnableGraphics=false, or EnableGraphics=true but zero overlays survived resolution, the label plumbing is untouched — the cut stage outputs directly to [vout] exactly as before Phase 3), then chains one drawbox (semi-transparent background, enable='between(t,start,end)', skipped entirely when OverlayBoxColor = "none") plus one drawtext per overlay (plus a second smaller drawtext when Subtext is non-empty), with the LAST overlay's final filter becoming the new [vout] that -map continues to reference. FilterComplexScriptThreshold's condition now also checks the built filter STRING LENGTH (spans.Count > 64 || filterComplex.Length > 4000), since overlays can make one long filter string even with very few cut segments.

Soft-failure discipline: the cut must never become hostage to graphics​

Every graphics-specific failure mode degrades to "no graphics applied", never to a failed compile — the cut is the primary deliverable:

SituationOutcome
GraphicsPlan configured but unresolvable/invalid JSONCut proceeds with no graphics; graphics.reason records why
GraphicsPlan not configured at all (null)Cut proceeds with no graphics; no error
An overlay's PlacementId not in OfferedPlacementIdsThat ONE overlay dropped (unknown_placement_id); the rest still apply
More overlays than MaxOverlaysExcess dropped (max_overlays_exceeded), in original order
Sanitized text ends up emptyThat overlay dropped (empty_text_after_sanitization)
Placement's window falls entirely in a cut gapThat overlay dropped (cut_away)
drawtext filter unavailable in this ffmpeg buildALL overlays skipped; graphics.unavailable = true
EnableGraphics=true with Mode=StreamCopyStep FAILS GRAPHICS_REQUIRE_REENCODE — the one graphics failure that IS a hard failure, since it is a pure config error caught before any resolution work, not a soft runtime condition

The one exception above (GRAPHICS_REQUIRE_REENCODE) is deliberate: drawtext/drawbox filters have no stream-copy equivalent, so this is a workflow-author config mistake to fix, not a runtime condition to degrade around.

VideoCompileStepExecutor.BuildEdl/the step's output summary JSON gain a graphics block — present ONLY when EnableGraphics=true (when false, the EDL/output shape is byte-identical to the pre-Phase-3 compile path):

{
"graphics": {
"enabled": true,
"applied": true,
"appliedOverlayCount": 1,
"droppedOverlays": [{ "placementId": "p7", "reason": "unknown_placement_id" }],
"unavailable": false
}
}

drawtext availability probe​

drawtext needs libfreetype (Alpine's font-dejavu package, which the WorkflowEngine Dockerfile now installs alongside ffmpeg) and a font file (VideoEditingOptions.FontFilePath, default /usr/share/fonts/dejavu/DejaVuSans.ttf, wired through VIDEO_FONT_FILE). A one-off runtime check (ffmpeg -hide_banner -filters, checking for drawtext in the output) is cached for the process lifetime — never re-probed per step. If unavailable (an unrebuilt image, or a dev environment missing the font package), ALL overlays are skipped and graphics.unavailable = true is recorded, but the cut video is still produced successfully — never fails an otherwise-successful encode over a missing font/filter.

Not built by Phase 3​

Text/box drawtext-drawbox overlays with fade in/out, OR a rendered Remotion asset overlay (see Rendered-asset overlays above) — Phase 3 does NOT apply Ken-Burns zoompan (Phase 1's KenBurnsCandidate remains identification-only), does NOT burn in subtitles, and does NOT do transitions between cuts. See Explicitly not built below, which is unchanged by this phase except for graphics moving out of "not built" and into this section.


Background music​

An optional background-music bed, mixed under the dialogue during VideoCompile's encode and ducked automatically during non-speech windows, planned by a fourth built-in agent that picks among uploaded tracks — the same structural shape Motion graphics (Phase 3) established: a deterministic candidate list from VideoAnalyze, an agent that chooses among opaque offered ids plus a handful of enum words, and VideoCompileStepExecutor alone resolving those words to real ffmpeg behavior. Off by default (VideoAnalyzeStepConfig.OfferMusicTracks = false, VideoCompileStepConfig.EnableMusic = false) — EnableMusic = false leaves the compile path byte-identical to the pre-music behavior.

Why a deterministic volume envelope, not sidechaincompress​

MusicMixPlanner/MusicMixFilterBuilder duck the music bed via a deterministic, keyframed volume=eval=frame envelope computed from the analysis artifact's own silence gaps/transcript segments — never a runtime audio-level compressor (sidechaincompress). Two reasons this codebase deliberately does not use a sidechain compressor here:

  1. The analysis artifact already carries silence gaps (available with zero dependency on ASR) and transcript segments — exactly the physically-grounded "speech has actually stopped" signal VideoCompileStepExecutor.ExtendSegmentEndTowardNextSilence already trusts over ASR boundaries — so there is no need to infer ducking windows from the waveform at encode time at all.
  2. A sidechain compressor's behavior depends on the actual waveform at encode time, so nothing could assert an exact filter string for it, explain "the music was lifted in these 3 windows" in the step's own output JSON, or guarantee sane behavior on a source whose dialogue track already has music baked in. The deterministic envelope, by contrast, is something MusicMixPlanner.PlanLiftWindows decides entirely in C# and MusicMixFilterBuilder.BuildVolumeExpression turns into an EXACT, assertable ffmpeg filter string.

Candidate discovery (VideoAnalyze)​

When OfferMusicTracks = true, VideoAnalyzeStepExecutor enumerates every audio/* project file as an m{n} music-track candidate (view.musicTracks), capped by MaxMusicTracks (default 20). This is project-level, not per-source — unlike every other candidate list this feature offers, one candidate list regardless of how many source clips the step analyzed. Candidates are never ffprobed here — the fit/duration policy is resolved server-side at compile time regardless of a candidate's exact length, so probing every candidate here would only cost N downloads for a list the agent picks at most one item from. A ListFilesAsync failure degrades to zero candidates; this never fails the step. view.musicTracks is shown as ONE ATOMIC ARRAY, deliberately not gated on VisualDetail (music has nothing to do with visual detail) — it is dropped as a whole array, after detail has already degraded all the way to None, before the per-item drop loop ever runs, mirroring the discipline Motion graphics (Phase 3) established for view.placements. "Offered" means exactly "the whole musicTracks array survived to the final view" — VideoAnalysisArtifact.OfferedMusicIds is empty whenever it was suppressed for budget.

m{n} id isolation​

Music-track ids (m{n}) are their own namespace, separate from shot/silence/segment ids (s{n}/g{n}/t{n}), placement ids (p{n}), and look-group ids (k{n}). OfferedMusicIds is its own separate list — a music-track id must never be validated against OfferedIds/ OfferedPlacementIds and vice versa — and, like placement/look-group ids, is deliberately NOT resolvable by VideoCompileStepExecutor.BuildIdTimeIndex: a Keep span naming an m{n} id fails UNKNOWN_ID exactly like any other id that index does not contain.

AgentType.MusicSupervisor and MusicPlanOutput​

An ordinary StepType.Agent step, following the exact precedent VideoStoryEditor/ MotionGraphicsPlanner set. Given the story editor's decision (or the same bounded view) plus view.musicTracks, it picks AT MOST ONE track plus a few coarse settings:

public class MusicPlanOutput
{
public string TrackId { get; set; } = ""; // must be in OfferedMusicIds; empty = no track suits the edit
public string Intensity { get; set; } = ""; // Quiet | Balanced | Feature — never a dB number
public string Ducking { get; set; } = ""; // Off | Light | Normal | Heavy — never a dB number
public string Fit { get; set; } = ""; // LoopToFit | PlayOnce
public string Reason { get; set; } = "";
public string PlanRationale { get; set; } = "";
}

The same rushcut invariant extended again: every property is a plain string, so there is no numeric/time-bearing CLR type to even ban — enforced by MusicPlanOutputInvariantTests. The model's only contributions are an opaque TrackId drawn from the set it was actually offered plus the three enum-word choices; VideoCompileStepExecutor alone resolves those to dB levels/ffmpeg behavior. This agent is entirely optional: VideoCompileStepConfig.MusicTrackProjectFileId, set directly by the workflow author, delivers the whole capability (a fixed track, default settings) without this agent at all. Same minimal read-only project-context + FailWorkflow tool scope as VideoStoryEditor — no render/sandbox escape hatch the way MotionGraphicsPlanner has, since there is no media for this agent to produce itself.

Resolution: ResolveMusicAsync, soft-failure throughout​

VideoCompileStepExecutor.ResolveMusicAsync runs only when EnableMusic = true, and — like every other stage in this feature — never fails the compile, only degrades to "no music applied":

SituationOutcome
MusicPlan configured but unresolvable, or not valid JSONDropped (plan_unresolved / plan_invalid_json); falls through to MusicTrackProjectFileId if set, else no music
Plan's TrackId not in OfferedMusicIds, or names no known candidateDropped (unknown_track_id); same fallthrough
Resolved project file not found in the project, or not an audio/* mime typeDropped (track_not_in_project / track_not_audio); no music
Track download or ffprobe fails, or reports zero duration/no audio streamDropped (track_download_or_probe_failed); no music
This ffmpeg build's amix filter has no normalize optionmusic.unavailable = true; no music (the cut still succeeds)

The amix normalize-option probe (IsAmixNormalizeAvailableAsync) is cached for the process lifetime, mirroring IsDrawtextAvailableAsync's pattern exactly — normalize=0 is load-bearing, not cosmetic: without it amix silently halves every input's level, including the dialogue track, so a missing option must degrade ALL music rather than risk quietly reducing dialogue loudness.

Resolution order for which track plays: MusicPlan (the agent's choice) first, then MusicTrackProjectFileId (the deterministic workflow-author-configured fallback) if the plan did not resolve to a usable track, else no music at all.

Enum words to ffmpeg behavior​

Intensity/Ducking/Fit are resolved entirely server-side, exactly mirroring how Motion graphics resolves Duration/Emphasis:

  • Intensity (Quiet/Balanced/Feature) → bed level in dBFS via MusicBedQuietDb/ MusicBedBalancedDb/MusicBedFeatureDb (defaults -26/-20/-14), clamped [-40, -6].
  • Ducking (Off/Light/Normal/Heavy) → attenuation below the bed while dialogue is present, via MusicDuckLightDb/MusicDuckNormalDb/MusicDuckHeavyDb (defaults -6/-11/ -18), clamped [-30, 0]. Off (from either the model or VideoCompileStepConfig.MusicDucking = MusicDuckingMode.Off) collapses lift-window planning entirely — a flat ducked bed throughout, no trapezoid expression.
  • Fit (LoopToFit/PlayOnce) → LoopToFit adds -stream_loop -1 to the music input and trims to the edit's exact frame-quantized length; PlayOnce trims to the track's own length when shorter than the edit, no loop.

MusicMixPlanner: where to lift the bed​

MusicMixPlanner.PlanLiftWindows (pure, no I/O, no ffmpeg — the OverlayPlacementBuilder precedent for this feature's other deterministic planning code) decides WHERE, on the compiled edit's own OUTPUT timeline, the bed should rise back toward its unducked level. Basis selection, in order — the first one with any data wins:

  1. silenceGaps (preferred) — the artifact's own detected silence spans, mapped through the same MapSourceWindowToOutput helper Motion graphics uses for placement windows.
  2. speechComplement — else, the gaps BETWEEN transcript segments, computed per source clip.
  3. noSpeechDetected — else, a single lift window spanning the WHOLE output: there is no dialogue anywhere in the kept edit, so the bed should not be needlessly ducked for the whole video.

Candidate windows are merged (closer together than MusicLiftMergeMs), dropped below MinMusicLiftWindowMs (and always below twice the gain ramp, so a lift too short to fully ramp never reads as pumping), and capped at MaxMusicLiftWindows (longest kept, re-sorted chronologically).

MusicMixFilterBuilder.BuildVolumeExpression turns the plan into an exact volume=eval=frame expression: zero lift windows collapse to the bare ducked-gain constant; one window is a single trapezoid (0 outside [start, end], ramping linearly to 1 across MusicDuckRampMs INSIDE each end of the window, so a lift is never above the ducked level exactly at a boundary); more than one window nests binary max(...) calls (ffmpeg's eval has no n-ary max). BuildMixStage's final amix uses normalize=0 (see above), duration=first (pins the mixed output's length to the DIALOGUE input, so a looped/infinite music input can never extend the file), and dropout_transition=0 (avoids a gain re-ramp when the music branch ends before the dialogue does, for PlayOnce with a track shorter than the edit).

Skipped entirely when the output has no dialogue audio​

When the compiled output has no dialogue audio at all — a single audio-less source clip, or (in a multi-source compile) a mix where NOT ONE referenced clip has an audio stream — there is nothing to duck against. VideoCompileStepExecutor computes hasDialogueAudioInOutput (isMultiSource ? anySourceHasAudio : sourceHasAudio) once, after source download/probe, and threads it into ResolveMusicAsync as hasDialogueAudio. That flag folds into the SAME duckingOff switch that already collapses lift-window planning to a flat, undocked bed level (duckingOff = configDuckingOff || !hasDialogueAudio) — so silence-gap/speech-complement ducking windows are never planned against dialogue that will not exist in the output. music.duckBasis records "no_dialogue_audio" explicitly in this case (distinct from "none", which means ducking was simply turned off by config/the model while dialogue audio does exist), and dialogueHeadroom reports {"applicable": false, "reason": "no_dialogue_audio_in_output"} instead of computing a headroom number against Phase 1 loudness data for audio that was dropped from the output entirely. The music bed itself is unaffected by this — a track can still be mixed in as the entire soundtrack of an otherwise-silent edit; only the speech-aware ducking behavior is skipped.

Narration is the exception. When the output has no dialogue audio but does carry voiceover lines, those lines ARE the speech to duck under: duckingOff stays false (unless config or the model turned ducking off), the bed — the LIFTED level — sits NarrationGapLiftDb (12 dB) above the intensity's bed (capped at -6 dB), and PlanLiftWindows runs with no silences or segments (noSpeechDetected, the whole program lifted) minus the voiceover windows, so the music plays full between lines and ducks by the ducking word under each one. music.duckBasis is "narration" and bedBasis is "narration_gaps". Balanced/Normal: -8 dB between lines, -19 dB under them, against the constant -20 dB the bed held before.

Review evidence: dialogueHeadroom​

VideoCompileStepExecutor.BuildDialogueHeadroom computes deterministic evidence for AgentType.VideoReviewAgent's StepType.ReviewLoop step: the duration-weighted mean dialogue RMS across the KEPT spans only (from Phase 1's per-shot VideoAnalysisShotAudio.RmsDbfs) against the resolved ducked-music level — a hard, server-computed headroom number, not something a model estimates from audio it cannot hear:

{ "applicable": true, "meanDialogueRmsDbfs": -22.4, "duckedMusicDbfs": -31.0, "headroomDb": 8.6 }

applicable: false when no kept shot carries a Phase 1 audio descriptor (AnalyzeAudioLevels was off/degraded) or, per the previous section, when the output has no dialogue audio at all.

EDL / output JSON shape​

Present only when EnableMusic = true (byte-identical to the pre-music compile path otherwise):

{
"music": {
"enabled": true, "applied": true, "unavailable": false,
"source": "plan", "trackId": "m1", "trackName": "ambient-bed.mp3",
"intensity": "Balanced", "ducking": "Normal", "fit": "LoopToFit",
"bedDbfs": -20, "duckedDbfs": -31,
"trackDurationSec": 42.0, "outputDurationSec": 96.3, "loops": 3,
"playEndSec": 96.3, "fadeInSec": 1.5, "fadeOutSec": 2.5,
"duckBasis": "silenceGaps", "liftWindows": 4, "liftCoveragePct": 18.2, "speechCoveragePct": 71.4,
"dialogueHeadroom": { "applicable": true, "meanDialogueRmsDbfs": -22.4, "duckedMusicDbfs": -31.0, "headroomDb": 8.6 },
"dropped": []
}
}

source is "none" / "plan" / "config" (which of MusicPlan/MusicTrackProjectFileId actually supplied the track); dropped is a list of {reason, trackId} entries recording every soft-failure the table above allows, empty when music applied cleanly.


Cutting to the beat​

A music-driven reel wants its cuts on the beat and its strongest shot on the drop. Three pieces make that possible without the editor ever naming a time.

Measuring the music (MusicBeatAnalyzer)​

Pure C#, no model, no new dependency (Services/Video/MusicBeatAnalyzer.cs). The track is decoded once through the existing IAudioExtractor (16 kHz mono 16-bit WAV, read by WavRmsSampler.ReadMonoSamples), then:

  1. Onset strength — the positive frame-to-frame flux of log energy (10 ms hop, 40 ms window, floored 60 dB under the loudest frame), with its ±0.25 s local mean removed. Each frame's flux is timed at its window's END, where the new energy entered.
  2. Tempo — the autocorrelation peak of the onset envelope between 70 and 180 BPM (parabolic interpolation between lags), then refined together with the beat phase by maximizing the mean onset strength sampled on the beat grid (±1.5 BPM in 0.01 BPM steps, one-hop phase steps). Confidence is how far that grid's onset mean stands above the envelope's own mean.
  3. Energy — the RMS level of every beat (each window opens a quarter beat early, so an attack a few milliseconds before the grid line belongs to its own beat). The drop is the beat whose following ~4 s outweighs the preceding ~4 s the most — the earliest such jump within 80% of the biggest, and at least 3 dB. The outro is the beat after the drop where everything that follows sits furthest (at least 4 dB) under the preceding ~4 s.
  4. Bars are four beats, aligned so the drop is a downbeat (with no drop, on the beat offset carrying the most onset strength).

Golden fixture: "The Last Point" (project file 226a445b-9957-4552-8c99-4736746bbd0b in the local stack). The numpy prototype it was cut with found 134.25 BPM, phase 0.115 s, drop 13.9 s, outro ~42.6 s over the reel's first 49 s. This analyzer measures 135.0 BPM (identically on the whole 124.6 s track and on 30/49/60 s prefixes), drop 13.93 s and, over the first 49 s, outro 42.8 s; it recovers a synthetic 134.25 BPM click track to within 0.3 BPM, so the 0.75 BPM gap is attributed to the prototype's coarser whole-lag envelope. MusicBeatAnalyzerTests holds synthetic click tracks (always run) and the real track (gated on REELBOLT_BEAT_FIXTURE).

Showing the editor (AnalyzeMusicBeats)​

VideoAnalyzeStepConfig.AnalyzeMusicBeats (default false, byte-identical) measures every offered music track up to MaxBeatAnalyzedTracks (6; each costs one download and decode) and the track named by MusicBeatTrackProjectFileId — the track the edit is cut to, usually the compile's MusicTrackProjectFileId. It is a field of THIS step rather than read from the compile step's config because the step cache keys a step on its own config only. Facts land in the artifact (VideoAnalysisMusicCandidate.Beat, VideoAnalysisArtifact.Music) and in the view as words and seconds (MusicBeatProbe.ViewNode):

"music": { "name": "the-last-point.mp3",
"beat": { "bpm": 135.0, "tempo": "fast", "beatSec": 0.444, "barSec": 1.778, "lengthSec": 124.6,
"grid": "steady", "drop": "bar 8, at 13.9 s", "dropSec": 13.93,
"outro": "bar 24, at 42.8 s", "outroSec": 42.83 } }

The VideoStoryEditor prompt (both copies) gained a separate section, "Cutting to the beat (only when the brief asks)": plan runs in whole bars, put the strongest moment right after the drop by letting the runs before it add up to its time, end around the outro — and still choose only ids: VideoEditDecisionOutput and its no-timestamp invariant are untouched.

Snapping the cut (BeatSync)​

VideoCompileStepConfig.BeatSync (VideoBeatSync: Off default and byte-identical, Beat, Bar) runs right after the cut list is final (padding, coalescing, caps, seam bridges) and before the output timeline is built, so everything downstream — music, narration, overlays, inserts, the EDL — sees the snapped program. It needs EnableMusic; the track is the music plan's offered choice, else MusicTrackProjectFileId, and its grid comes from the analyze step's measurement when there is one (gridSource: "analysis"), else it is decoded and measured here ("measured"). Background music starts at program time 0 and loops to fit, so the grid is the track's own, repeated per loop (BeatSyncPlanner.LoopedGrid).

BeatSyncPlanner.Plan walks the spans in program order. Each span's cut moves EARLIER, to the last grid line (bars, or beats) that still leaves the span at least one beat long — only the TAIL of a span is trimmed, never extended past its own footage. Bar falls back to a beat when no bar line fits; with neither, the span is kept whole and the next cut re-aligns. When the drop is at least two bars in and falls inside a span at least a beat after its start, that span is cut exactly on the drop, so the next shot lands with it. Contiguous kept shots were already merged into one span, so only real cuts move. Frame quantization afterwards moves a cut by at most one frame; an overlapping seam transition (TransitionPolicy) shifts later cuts by its overlap, so pair beat sync with hard cuts for exact alignment.

Narration is never cut short by a snap. With EnableVoiceover, the planner first reads the Voiceover step's own output (VideoCompileStepExecutor.NarrationSpanFloors: each ok line's anchorId and measured durationSec) and gives every span a floor — from the span's start to the end of the last narration line anchored inside it (a line starts at its anchor or right after the previous line in that span, whichever is later), plus 0.25 s. A snapped cut must land at or after that floor; when no grid line fits before the span's own end, the span is kept whole. Found live: a narrated reel cut on the bar trimmed 3.75 s of tails and 5 of its 10 lines ran onto the next shot (voiceover.fit.overflowLineCount 5, worst 1.47 s). beatSync.narrationProtectedSpans counts the spans that carried a floor.

The output summary and the EDL carry, only when BeatSync is not Off:

"beatSync": { "mode": "Bar", "applied": true, "track": "the-last-point.mp3", "gridSource": "analysis",
"bpm": 135.0, "beatSec": 0.4445, "barSec": 1.7779, "dropSec": 13.934, "dropAligned": true,
"keptWholeSpans": 0, "trimmedSec": 1.42,
"cuts": [ { "segment": 0, "outputEndSec": 3.2, "snappedEndSec": 3.044, "trimmedSec": 0.156, "grid": "bar" } ] }

or applied: false with a reason (music_not_enabled, no_music_track, track_not_in_project, track_not_audio, no_beat_grid, beat_sync_failed) and the cut exactly as the editor made it.


Seam transitions and the program envelope​

VideoCompileStepExecutor can apply a short transition treatment at each cut-seam between two kept spans, and a fade at the very start/end of the whole compiled program — both entirely deterministic, selected from the same measured shot/seam evidence surfaced to VideoReviewAgent (see Review evidence below), never from an agent. No step in this pipeline can request, add, remove, lengthen, or shorten a transition: VideoStoryEditorAgent's prompt states outright it "cannot create, request, or describe a transition, fade, dissolve, or effect of any kind," and VideoReviewAgentImpl's prompt is told the same choice is made "deterministically by the compile step" and to never phrase a fix as "add a fade" or "soften that cut" — only ever as a different choice of WHICH SPANS to keep. This is the same "rule table, not a model" discipline Motion graphics's Duration/Emphasis words and Background music's Intensity/Ducking/Fit words already established for this feature — except here there is no agent step in the loop at all, deterministic end to end.

Inter-cut transitions vs. the program envelope​

Two independent things, controlled by separate config, and easy to conflate:

  • Inter-cut transitions happen at every internal seam between two kept spans — a hard cut, a brief audio-only declick ramp (AudioSeamRampMs), a short crossfade dissolve (DissolveMs), or a dip-to-black/dip-cut (DipToBlackMs/DipCutMs) — chosen per seam from that seam's own measured properties, capped by MaxTransitionMs/MaxTransitionRatioPct so a transition can never eat a meaningful fraction of either neighbouring segment.
  • The program envelope is the single fade-in at the very start and fade-out at the very end of the WHOLE compiled file — video (ProgramFadeInMs/ProgramFadeOutMs) and audio (ProgramAudioFadeInMs/ProgramAudioFadeOutMs) tracked separately, since a video dip-to-black and an audio fade need not move in lockstep. This has nothing to do with any internal seam; it exists purely so a finished piece never starts or ends on a hard, un-eased frame — the same weight VideoStoryEditorAgent's own "The opening and the closing" prompt section already places on the first and last kept span, applied here at the encode instead of the edit-decision stage. The audio envelope never applies to narration: voiceover lines are mixed after it, and a last line running past the picture holds the final frame until it has finished — see voiceover.md "Narration and the end of the program".

TransitionPolicy​

VideoCompileStepConfig.TransitionPolicy (a VideoTransitionPolicy enum) gates the whole inter-cut system: at its most permissive setting — "Auto", what all three video-derush-edit* templates now request — the rule table is free to apply whichever treatment a seam's own measurements call for; turned off entirely, every seam falls back to a hard cut and the compile path is byte-identical to the pre-transition behavior, the same "off by default, byte-identical without it" guarantee EnableGraphics/EnableMusic already give this feature. The full set of intermediate levels the enum exposes belongs to the sibling change that introduces VideoTransitionPolicy itself; the contract fixed here, and depended on by every other piece of this feature (the prompts above included), is that every policy level except Editor selects a treatment FROM measured data, and Editor lets the editor choose only by WORD — see Editor-chosen transitions.

Editor-chosen transitions​

TransitionPolicy = Editor (appended last) hands each seam to the editor agent. Every VideoEditKeepSpan may carry Transition — a preset WORD — and TransitionSpeed — Snap, Quick, Smooth or Slow — naming how that span comes IN; both are strings, so the rushcut invariant (no number in the decision) holds, and PickupPlanOutputInvariantTests pins the widened property set. EditorTransitionCatalog alone resolves them: 56 presets map one-to-one onto ffmpeg's built-in xfade transitions (Fade, FlashWhite, Zoom, Pixelize, SlideLeft, PushUp, RevealRight, WipeDown, SmoothLeft, CircleOpen, BarnDoorOpen, SliceLeft, WindRight, SqueezeHorizontal, …; a raw xfade name is accepted too), DipToBlack takes the DipToBlack treatment and every other preset the Dissolve one with its own xfade name, Cut is a hard cut, and the speeds are 0.2 / 0.35 / 0.6 / 1.0 s (Quick when a preset names none).

A resolved span takes the choice of the kept span it starts with (ResolvedSpan.OriginKeepFromIds[0]); a seam whose span names nothing, or names an unknown preset (reported in transitions.editor.unknownPresets), takes the Auto rule table. An editor choice is NOT exempt from the safety passes: no xfade demotes it to DipCut, a seam touching a chroma-plate span is still demoted, and the per-seam (half the shorter span, MaxTransitionMs) and neighbour-sum (60%) overlap clamps still apply. It IS exempt from the density cap (MaxTransitionRatioPct), which only downgrades rule-table seams: an editor that put a transition there meant it. Its seams carry the rule E:<Preset>:<Speed>, and transitions.editor reports chosenCount and every chosen seam (seam, choice, treatment, transition, durationSec). The VideoStoryEditor prompt (both copies) lists every preset and teaches restraint — most seams are cuts, directional moves keep one direction, hard hitters go on a beat. Under any other policy the two fields are ignored.

SectionBreakGapMs and silence-anchored seams​

A seam that falls on a long-enough silence gap (SectionBreakGapMs) reads as a genuine section break rather than an ordinary mid-sentence cut, and the rule table treats it differently from a seam with no silence on either side — the same kind of distinction seamCheck already exposes to the reviewer via cutOutMotion/cutInMotion for motion, applied here to silence instead.

MaxTransitionSegments: the filtergraph-buffering guardrail​

A segmented single-source encode already writes one select/aselect filtergraph entry per kept segment (see MaxSegments above); a crossfade-style transition needs to buffer and blend TWO adjacent segments at once instead of switching between them instantaneously, which multiplies the filtergraph's memory/CPU cost per transition rather than per segment. MaxTransitionSegments caps how many segments a compile is willing to apply transitions across at all — above the cap, transitions are skipped for the WHOLE compile (every seam falls back to a hard cut) rather than risking an ffmpeg process that OOMs or times out on a long, heavily-cut edit. This mirrors MaxSegments's own "switch to a scratch-file filtergraph above ~64 segments" guardrail in spirit: a structural limit protecting the ffmpeg process, not a quality knob.

Review evidence​

VideoCompileStepExecutor's output JSON surfaces transitions (policy, appliedCount, and a treatments breakdown) and programFade alongside the existing sentenceCheck/graphics/music blocks, plus three more deterministic checks that arrived alongside the transition system: openingCheck (the sentenceCheck mirror for the FIRST kept span instead of the last), seamCheck (every inter-cut seam's measured look/motion continuity, both an aggregate count and an itemized list of the notable ones), and pacing (segmentCount/meanSegmentSec/ medianSegmentSec/shortSegmentPct). All of these are computed once per compile and read — never re-derived — by AgentType.VideoReviewAgent's ReviewLoop step exactly like sentenceCheck/ graphics/music already are; see VideoReviewAgentImpl's prompt for the exact score caps and remediation rules each one drives. When the transitions node is absent entirely, transitions were switched off for that workflow and the reviewer is told explicitly not to penalize hard cuts.


Semantic visual dimensions (Phase 4)​

Seven deterministic dimensions (D1-D7) added on top of Phase 1's scene/visual analysis, plus prompt/contact-sheet/persistence changes to Phase 2's vision captioning. D1-D4 and D6 are free — derived from the same low-res grid data Phase 1 already samples, so they default on. D5 (audio character) rides AnalyzeAudioLevels the same way. D7 (backlit) is a byproduct of D6's region data, also free. Only sharpness (a D-adjacent dimension, not part of D1-D7's own numbering) costs a genuinely new ffmpeg invocation per measured shot, so it alone stays opt-in.

Every new stage below follows Phase 1's original discipline: degrade, never fail the step. A classification that cannot be computed (near-monochrome frame, too few audio windows, no clean letterbox bars) simply omits that field or id — it never throws out of the analysis pass, and it never blocks silence/shot detection, transcription, or any other deterministic stage from completing.

#DimensionFormula (informal)GateDocumented limitation
D1Color temperature (Warmth/Tint/ColorTemperatureClass)Warmth = clamp((meanR − meanB) / 0.25, −1, 1); Tint is the same shape against G vs. (R+B)/2. ColorTemperatureClass is Warm/Cool at |Warmth| ≥ 0.20, else Neutral; a near-monochrome frame (SaturationMean < 0.05) is always Neutral regardless of WarmthAnalyzeColorGradingA single dominant colored object (not the lighting) can skew the whole-frame mean; this is a frame-average heuristic, not a white-balance measurement
D2Tone curve (BlackPoint/WhitePoint/ToneClass)5th/95th percentile of the luma histogram (nearest-rank). ToneClass order is load-bearing — an actual exposure defect always outranks a stylistic read: Blown (clipped highlight ratio > 5%) → Crushed (crushed black ratio > 5%) → Flat (dynamic range < 0.45 and black point > 0.10 — lifted blacks + compressed range, i.e. log/ungraded) → Contrasty (dynamic range > 0.75 and black point < 0.06) → NormalAnalyzeColorGradingFlat on a whole look group is a property of the SOURCE footage (ungraded log), not a per-shot defect — both agent prompts say so explicitly
D3Saturation character (SaturationClass)Pure threshold projection of the pre-existing SaturationMean: Muted (< 0.18) / Natural / Vivid (> 0.42)AnalyzeColorGradingSame mean-based coarseness as D1
D4Look grouping (view.lookGroups, ids k{n})Six-float LookSignature (Warmth, Tint, BrightnessMean, BlackPoint, WhitePoint, SaturationMean) per shot; LookDistance is a weighted L1 distance normalized per-component to 0..1 (weights 0.30/0.10/0.25/0.15/0.10/0.10, summing to 1.0 so Similarity = 1 − Distance lands in 0..1); single-linkage clustering over all pairs (not windowed like near-duplicate grouping, since a shared look deliberately links non-adjacent shots/clips) at LookSimilarityThreshold (default 0.88). LookRank orders group members by ascending distance to the group centroid — rank 0 is the most representative shotDetectLookGroupsO(shots²) — trivially cheap at the shot counts this feature targets, but would need revisiting at extreme shot counts
D5Audio character (char under "a")ShotAudioAnalyzer.Analyze: crest factor (peak − RMS dB), 10th-percentile window RMS as a noise floor, level stability (1 − clamp(stdDev/12dB, 0, 1)), and zero-crossing rate. Classification order is load-bearing (Dialogue is the fallthrough, never a positive claim): Silent (RMS ≤ −50 dBFS) → Music (stable + compressed + low ZCR) → Noisy (high noise floor, low speech ratio) → Ambient (quiet, low speech ratio) → DialogueAnalyzeAudioLevelsDeliberately a ZCR/crest/noise-floor heuristic, not a spectral (FFT) classifier — this repo's only audio test fixtures are synthetic sine tones, and thresholds tuned against a 440 Hz sine would pass CI while misclassifying real footage (judgment call 7)
D6Letterbox/pillarbox (ActiveCrop)Scans the already-materialized luma frames for rows/columns whose luma stays below a tolerant black threshold (16/255, tolerant of compression noise inside a true matte) across every sampled frame, capped at 40% of the frame dimension (beyond that it reads as a dark scene, not bars)DetectLetterbox (default true — free, no second luma pass)Under-reports soft/gradient letterbox edges — the threshold expects a clean black bar, not a feathered one
D7Backlit candidate (BacklitCandidate)Byproduct of D6's region data: the center region (R4) is markedly darker than the average of the surrounding border regions (borderLuma − centerLuma > 0.18) while itself being dark (centerLuma < 0.35)Free whenever regions are computedNamed and surfaced as a candidate, not a claim (mirrors KenBurnsCandidate's precedent) — fed to the vision prompt so a model that can actually see the frame turns it into (or rejects) an actual assessment; Full detail only
—Sharpness (Sharpness, D-adjacent)Native-resolution, square, centered grayscale patch (FfmpegArgvBuilder.BuildSharpnessPatchArgs, deliberately its own ffmpeg invocation so a fault there can never take Phase 2 keyframe extraction down with it) → discrete 4-neighbour Laplacian → variance over the interior, normalized against a documented heuristic constant. Only ever compared BETWEEN shots of the same source, never as an absolute unitDetectSharpness (default false — one extra ffmpeg call per measured shot)Costs real wall-clock/CPU unlike D1-D4/D6/D7, which is why it alone stays opt-in; capped by MaxSharpnessShots, a genuinely step-wide budget (like MaxCaptionedShots) computed AFTER the per-shot analyze loop and BEFORE that source's own duplicate grouping, so a real measured value (when available) — not the neutral placeholder — reaches ComputeTakeQuality's best-take scoring

k{n} isolation​

Look-group ids (k{n}) are a purely descriptive namespace, structurally different from every other id this feature offers a model. Shot/silence/segment ids (s{n}/g{n}/t{n}) are offered to VideoStoryEditorAgent and resolved by VideoCompileStepExecutor.BuildIdTimeIndex; placement ids (p{n}) and music-track ids (m{n}) are separate offered namespaces resolved by their own executor paths. k{n} is never offered to any agent at all — there is no OfferedLookIds list, deliberately — and no structured output in this feature (VideoEditDecisionOutput, MotionGraphicsPlanOutput, MusicPlanOutput) has a field that can name a look group. A Keep span naming k{n} can therefore only be a hallucination, and BuildIdTimeIndex deliberately excludes k{n} from its index so such a span fails UNKNOWN_ID exactly like any other id it does not contain — the same fate as a p{n}/m{n} id named in a Keep span.

View/artifact shape​

{
"view": {
"shots": [
{
"id": "s0", "startSec": 0.0, "endSec": 4.2, "durationSec": 4.2, "src": 0,
"v": {
"motion": 12, "move": "Static", "cutIn": "still", "cutOut": "still",
"bright": 58, "colors": ["#3a2c1e", "#c9a876"],
"temp": "Warm", "tone": "Normal", "sat": "Natural", "look": "k0", "crop": [0.0, 0.11, 1.0, 0.78]
},
"a": { "rms": 42, "speech": 71, "char": "Dialogue" },
"c": {
"summary": "A presenter gestures at a whiteboard while explaining a diagram.",
"scale": "Medium", "mood": "Focused", "tags": ["presenter", "whiteboard"],
"style": "flat ungraded log", "issues": []
}
}
],
"lookGroups": [
{ "id": "k0", "shotIds": ["s0", "s1", "s4"], "repShotId": "s0", "cohesion": 91, "temp": "Warm", "tone": "Normal", "sat": "Natural" }
]
},
"meta": {
"look": { "applied": true, "uniform": false, "groupCount": 1 },
"vision": { "mode": "Optional", "applied": true, "captionedShots": 6, "failedShots": 0, "persistedKeyframes": 0 }
}
}

"temp"/"tone"/"sat" and "look" live under a shot's existing "v" node — gated on AnalyzeColorGrading/DetectLookGroups and VisualDetail exactly like every other Phase 1 field, same degrade-before-drop discipline. "crop" (D6) appears only when a crop was actually detected — omitted, not null, when the frame is full-bleed. "char" (D5) is always present on the "a" node whenever audio levels were analyzed, including for the common Dialogue value — unlike the visual fields, there is no "absent means nothing to report" reading for D5, since every shot has SOME audio character. view.lookGroups is capped by MaxViewLookGroups; when the whole analysis is one uniform look, every per-shot "look" id is suppressed and meta.look.uniform is true instead — repeating the same group id on every shot would add bytes without adding information.

Progress weighting​

VideoAnalyzeProgressPlan's per-source stage list gained SampleSharpness (weight 4, only counted when DetectSharpness is on) between AnalyzeShots and GroupDuplicates, and the step-level list gained MatchLooks (weight 2, only counted when DetectLookGroups is on) between GroupDuplicates/multi-source join and ListMusicCandidates:

StageWeightLevel
SampleFrameGrid12Per-source
AnalyzeShots10Per-source
SampleSharpness4Per-source (only when DetectSharpness)
GroupDuplicates2Per-source
MatchLooks2Step-level (only when DetectLookGroups)
CaptionShots14Step-level

A disabled stage contributes zero weight and is skipped entirely rather than reported as an instant 0%-to-100% jump — the same "only enabled stages count toward the total" rule the original plan established for every other optional stage, so turning a Phase 4 dimension off never distorts the percentages reported for the stages that stayed on.

Vision-phase changes​

See Prompt priming, contact sheets, and persisted keyframes (Phase 4) under Vision captioning, and the multi-source hoist fix noted at the top of Executor wiring — both are Phase 4 changes to the Phase 2 captioning path, documented alongside Phase 2 rather than duplicated here.


Tracked screen inserts (Phase 5)​

Compositing a Remotion-rendered scene INTO a moving region of the source footage — the canonical case: a commercial shot with a phone held in frame against a green screen, where an app UI (rendered as its own Remotion composition) is inserted into the phone's screen area and moves and warps with the phone as the hand moves, not as a static overlay. Off by default at both ends (VideoAnalyzeStepConfig.DetectInsertRegions = false, VideoCompileStepConfig.EnableInserts = false); EnableInserts = false leaves the compile path byte-identical to before this phase.

The invariant this phase exists to protect: a motion-tracking transform is inherently per-frame numeric data — positions, corner coordinates, times. Every one of those numbers is computed by deterministic C# and consumed by deterministic C#. The model's entire contribution is (1) an opaque r{n} region id drawn from the set it was actually offered (VideoAnalysisArtifact.OfferedInsertRegionIds — a separate id namespace, never resolvable by BuildIdTimeIndex, "offered is stricter than exists" like every other id family) and (2) a rendered asset it produced itself via a real RenderVideoAndUploadToStorage call, validated against this execution's own outputFiles prefix — the exact RenderedAssetStorageKey precedent. ScreenInsert (on MotionGraphicsPlanOutput) has exactly three string properties — RegionId, RenderedAssetStorageKey, Reason — locked by MotionGraphicsPlanOutputInvariantTests, which also pins the exact property set so even a new string property is a visible, reviewed decision.

The tracking decision: chroma-plate quads, not general feature tracking​

ChromaQuadTracker (WorkflowEngine/Services/Video/ChromaQuadTracker.cs) is a pure, unit-tested static class. It runs over a dedicated, medium-resolution raw-RGB grid pass (the same IFrameGridSampler machinery Phase 1 uses, at ~320px on the long axis / 10fps instead of 32x18/2fps) and, per frame: builds a chroma mask by relative channel dominance (brightness-robust — g*10 > r*13+100 etc., for green/blue/magenta), finds the largest 4-connected component (BFS), grows it through a second, permissive mask (see "Dual-threshold chroma masking" below), fits a quadrilateral via the extreme-point method (TL=min(x+y), BR=max(x+y), TR=max(x−y), BL=min(x−y)), and gates on area ratio, component-vs-quad fill ratio (rejecting L-shapes and scattered noise) and quad degeneracy. Per-frame quads assemble into tracks (dropout gaps ≤ 3 frames bridged by linear interpolation; larger gaps or implausible centroid jumps split the track), corners get a small centered moving-average smooth, and each track carries a confidence plus qualitative size/motion/aspect descriptors.

The tracking grid preserves the source's aspect ratio​

The grid's dimensions are derived from the source's own probed aspect (DeriveGridSize), not fixed. InsertGridWidth/InsertGridHeight default to 0 = auto; a positive pair is an explicit override and is honoured verbatim.

This matters twice over. A fixed 320x180 grid squashes a 1080x1920 portrait source vertically by 10.7x, so (a) vertical corner precision collapses to ±10.7 source px against ±3.4 horizontal — past InsertOverscan's 0.02 default, which on real footage left a visible green fringe (measured: 0.773% residual green at the fixed grid versus 0.090% at an aspect-matched 202x360), and (b) every physical distance measured in grid space is wrong by the aspect mismatch. The most visible casualty was MeanAspectRatio, the one real number offered to MotionGraphicsPlanner: a phone plate whose true aspect is 0.47 was reported as 1.46 — telling the agent a portrait screen was landscape.

Auto-derivation fixes the precision half. The aspect half is fixed independently and belongs to defense in depth: Track takes the source's real sourceWidth/sourceHeight as explicit parameters used only for physical-distance measurement, so the reported aspect stays correct even if a workflow author pins a grid that disagrees with the source. 16:9 sources still derive exactly 320x180, so nothing changes for the common case.

Dual-threshold chroma masking, and what confidence actually measures​

The channel-dominance ratio test is brightness-robust; the absolute floor beside it is not. A single floor of g >= 60 sat in the middle of a real dim laptop panel's own brightness distribution — that plate's green channel measured 28..88 across the same physical screen — so the mask captured only 66% of it (8.97% of frame area against a tolerant test's 13.56%), and the extreme-point fit then fitted a confident, tidy quad to the brighter fragment, leaving a large triangular strip of raw green exposed in the composite.

Detection is therefore dual-threshold (hysteresis), the standard fix for exactly this:

  • the core mask (floor 60, unchanged margins) seeds detection with pixels that are unambiguously the plate;
  • the grow mask (floor 24 — where 8-bit channel ratios stop being meaningful at all — with proportionally relaxed margins) extends that seed, and only that seed, by 4-connectivity.

Growth can never jump to a different object: an unrelated green-ish blob is admitted only if it is physically contiguous with the high-confidence core, and if growth ever did merge scenery the existing MinFillRatio gate rejects the resulting non-quadrilateral blob exactly as before.

Confidence was coverage × mean fill ratio, with fill clamped to 1.0. That made the second term a constant in practice (a discretized quad fit normally over-fills slightly, ~1.02–1.07), so the score was effectively just coverage — and nothing in it could see under-detection, because every term was computed against what the strict mask happened to find. That is how a visibly wrong detection reported 0.9996 and sailed past MinInsertConfidence. It is now a product of four independent axes (the photometric and temporal terms were revised later — see Self-tuned growth and honest confidence):

TermAxisCatches
coveragetemporalframes where nothing was detected at all
meanFillScoregeometrica fit that does not bound its own component. Two-sided (fill or 1/fill), so a degenerate fit scores low instead of being clamped up to perfect — a 45° plate collapses two extreme points onto the same pixel and fits a triangle at fill ≈ 2.03, which the old clamp rewrote as 1.0
meanEdgeContrastphotometricdoes each fitted edge sit on a real plate boundary — chroma purity just inside the edge minus just outside it. Replaced meanChromaMargin (|core| / |grown|), which docked a correct dim-plate fit for needing growth and could not tell a fit that leaked into scenery from a good one
consistencytemporalthe fraction of frames whose four corners agree with their neighbours' temporal median — a physical plate moves continuously, a fit that flickers onto scenery does not

A quad with two coincident corners is now rejected outright rather than warped. On the real footage the net effect is: the laptop plate is detected at full extent (13.6% of frame, corners landing on the screen edge instead of ~100px inside it) and scored 0.63–0.64 instead of 0.9996, while the clean phone plate still scores 0.91–0.92. Note the goal is not to reject the laptop shot — with the mask fixed its detection is correct — but to make the score mean something, so raising MinInsertConfidence can actually exclude marginal plates.

Why marker/chroma-based, and what was rejected. General markerless planar tracking (KLT/feature-correspondence homography estimation) was evaluated and rejected for v1, in the same cost-benefit style as the "why ffmpeg is not in the sandbox" decision:

  • OpenCV via OpenCvSharp would be the honest way to do markerless tracking, but it is a large native dependency with no musl/Alpine binaries the WorkflowEngine image could consume without building OpenCV from source — hundreds of MB of image growth, a large native attack surface parsing untrusted media (the same class of risk the existing ffmpeg-hardening flags exist for), and a build pipeline burden, for a capability whose robustness could not be validated here against real footage anyway.
  • A from-scratch C# feature tracker could not honestly be claimed robust — pyramidal Lucas-Kanade plus RANSAC homography fitting is a genuine CV subsystem, not a helper class.
  • A chroma plate is the one target pure pixel statistics detect reliably — and it matches how this shot is actually produced in practice (a phone screen displaying solid green IS the standard on-set practice for exactly this composite). The constraint is stated plainly: the source footage must contain a uniform-color plate (green by default; blue/magenta configurable). Footage without one gets zero offered regions, and the agent is prompted to plan zero inserts.

The tracker's precision is bounded by the tracking grid (~0.3% of each frame dimension at the auto-derived default — ±6px at 1080p, now equal in both axes because the grid matches the source's aspect), softened by corner smoothing and by InsertOverscan expanding the insert slightly past the plate's edge. Raise InsertGridWidth/InsertGridHeight when tighter registration matters more than the larger grid buffer — but set BOTH, since a half-set pair falls back to auto.

The full numeric track (VideoInsertRegionTrack.Keyframes — per-sampled-frame normalized corner quads on the source timeline) lives ONLY in the analysis artifact. The bounded view offers view.insertRegions: {id, shotId, startSec, endSec, conf, size, motion, aspect, color} (+src when multi-source) — the window is READ-ONLY input exactly like view.placements' startSec/endSec, conf/size are bucketed words, and aspect is the one number the agent genuinely needs as input (to render suitably-proportioned content). The array follows the musicTracks budget discipline (unconditional, atomically suppressible — deliberately NOT the detail-gated placements discipline, since insert regions come from their own grid pass and are independent of Phase 1 visual analysis).

Agent surface​

No new agent: AgentType.MotionGraphicsPlanner gained an Inserts list on its existing MotionGraphicsPlanOutput (additive — every pre-existing plan deserializes with zero inserts), and its prompt (fallback + seeded, verbatim-locked by VideoStoryEditorPromptConsistencyTests) gained a "Tracked screen inserts" section.

Two things this addition originally missed, both found by running it. AgentType.MotionGraphicsDirector — the agent that performs the graphics room's structured synthesis, i.e. the one that actually emits a MotionGraphicsPlanOutput — never got that section, and its prompt closes by enumerating what to output as "an overlays list ... and a planRationale". Handed two real tracked plates it rendered both insert scenes, described them in its planRationale, and emitted an empty inserts list. Separately, DatabaseSeeder.GenerateMotionGraphicsPlanSchema — the hand-written schema documentation seeded onto every agent row — was never updated either, so the platform's own schema viewer described a two-field plan; and because existing rows only had OutputSchemaJson written when the column was empty, correcting the literal could not have reached an already-seeded deployment. All three are fixed, and SeededOutputSchemaDriftGuardTests now compares every seeded schema's top-level property names against its CLR output type so the documentation cannot silently fall behind the type again. The agent renders the insert content itself through the same sandbox+Remotion pipeline it already uses for rendered-asset overlays — but OPAQUE (a normal mp4, no alpha), since the whole rectangular frame is warped to fill the plate.

Compile: the corner-pin recipe​

VideoCompileStepExecutor.ResolveInsertsAsync (soft-failure throughout) validates each insert, maps the track's source window through the cut (OutputTimeline.MapWindowToOutput — clipped to the FIRST kept portion, like overlays), converts each surviving tracked keyframe to an output-frame-indexed pixel quad (BuildInsertKeyframes: per-keyframe MapToOutputSec, uniform downsample to MaxInsertExprKeyframes, centroid overscan expansion), downloads + ffprobe-validates the asset, and hands ScreenInsertFilterBuilder a fully-resolved numeric description. The filtergraph per insert — validated end-to-end against a real ffmpeg run during design:

[N:v] scale=(W-2)x(H-2), fps=canonical, pad to WxH with a 1px black border,
perspective sense=destination eval=frame (corner exprs piecewise-linear in `in`),
setpts +outputStart -> warped content
color=white (W-2)x(H-2), pad 1px black border, format=gray,
the SAME perspective exprs, the same setpts -> warped mask
alphamerge(content, mask) ; overlay at 0:0, enable='between(t,start,end)'

The 1px border is load-bearing: perspective edge-clamps out-of-range source coordinates, so an unbordered warp smears content across the whole frame outside the quad; bordering both the content and an all-white mask makes everything outside the warped quad black, which alphamerge turns into transparency (perspective itself supports no alpha format — that is why the mask branch exists). The insert stage renders BEFORE text/asset overlays (screen content is scene content; lower-thirds paint on top), its inputs sit between the asset-overlay inputs and the music input (preserving both existing index mappings), and corner expressions are piecewise-linear if(lt(in,f),a+(in-f0)*s,...) chains over the filter's per-frame in variable — every literal through FfmpegArgvFormat.Number. Not one model-originated character reaches the filter string. in is 1-based (ffmpeg evaluates it as the frame count plus one), so keyframe k is written at in = k + 1; until that was known every insert ran one frame ahead of its plate — see Tracking quality on the customer clips.

Soft-failure table​

SituationOutcome
EnableInserts=true with Mode=StreamCopyStep FAILS INSERTS_REQUIRE_REENCODE — pure config error, mirrors GRAPHICS_REQUIRE_REENCODE
No GraphicsPlan configured / plan unresolvable / invalid JSONNo inserts applied; inserts.reason records why
RegionId not in OfferedInsertRegionIdsThat insert dropped (unknown_region_id)
Same region chosen twiceSecond dropped (duplicate_region_id)
More inserts than MaxInsertsExcess dropped (max_inserts_exceeded)
Track confidence below MinInsertConfidenceDropped (confidence_below_threshold)
Asset key missing or outside this execution's outputFiles prefixDropped (invalid_asset_storage_key)
Asset download/probe failsDropped (asset_download_or_probe_failed)
Track window entirely inside a cut gap, or its own source clip contributed no output time at allDropped (cut_away)
perspective/alphamerge missing from the ffmpeg buildALL inserts skipped; inserts.unavailable = true
Window falls entirely inside a crossfade seam's blend regionDropped (window_consumed_by_transition_overlap)

The EDL/output summary gains an inserts block (present only when EnableInserts=true): {enabled, applied, appliedInsertCount, droppedInserts: [{regionId, reason}], unavailable}.

v1 scope, stated plainly​

What works: a chroma plate (green/blue/magenta) tracked as a deforming quadrilateral — translation, scale, rotation, and perspective skew all follow the plate, since all four corners are tracked independently and the warp re-evaluates per frame. On both encode paths: the original single-source select path and the segmented concat path used for multi-source compiles and crossfade transitions.

Inserts on multi-source compiles​

Inserts were originally skipped outright whenever the compile routed to the segmented encode, which meant a green-screen clip could not be used in any multi-source edit — the raw flat plate simply appeared in the final cut. That was encode-path plumbing, not a limit of the math: the corner-pin never depended on how the base video was assembled.

OutputTimeline is already source-aware, and both mapping calls pass the tracked region's own track.SourceIndex — MapWindowToOutput for the on-screen window, MapToOutputSec per corner keyframe. Both only ever intersect spans belonging to that region's own clip, so an insert resolves onto exactly the output time that clip contributed. Inertness during other clips' segments is not a special case: the composite is one overlay gated by enable='between(t,start,end)' over that window, so during output time cut from a different clip the overlay contributes nothing. This is the same mechanism OverlayAssetFilterBuilder already uses to gate an asset overlay to its own on-screen window.

The segmented path's stage order now mirrors the single-source path exactly — concat → colour grade → screen inserts → text overlays → asset overlays — and insert asset inputs sit after every source input and every asset-overlay input, before the music and SFX inputs (all downstream index offsets shifted accordingly).

Geometry needs the letterbox, not just the timeline. A tracked quad's corners are normalized to its OWN source frame. On the single-source path that is the canvas, so normalized corners times canvas dimensions is correct. On the segmented path with several clips every span is normalized with scale=cw:ch:force_original_aspect_ratio=decrease plus a centered pad, so a clip whose aspect differs from the canvas is letterboxed or pillarboxed and occupies only part of it. SourceCanvasFit reproduces that same arithmetic per source, and BuildInsertKeyframes maps through it. Without this the insert composites in the right place in time and the wrong place in space: caught on a real three-source compile where a 1080x1920 phone clip pillarboxed into the canvas had its insert drawn roughly three times too wide, spilling across the entire frame. Note the canvas is the FIRST KEPT span's source, not a fixed 1920x1080 — reorder the edit and the canvas changes with it, which is exactly why the fit is computed per source rather than assumed.

Overscan still expands about the plate's centroid in the source's own normalized space before the fit maps it, so InsertOverscan keeps meaning "a fraction of the plate" rather than "a fraction of the canvas". The mapping is affine, so the order is equivalent.

One case genuinely does not admit an insert: an overlapping seam treatment (dissolve/dip/whip) blends two spans into the same output frames, and a quad pinned to one clip's tracked geometry must not paint over a blended frame. OutputTimeline.ClipAwayTransitionOverlaps trims an insert's window back to its longest clean stretch, and drops it (window_consumed_by_transition_overlap) only if nothing survives. This is a no-op for every compile at the default TransitionPolicy = Off.

Deferred, deliberately: markerless tracking of arbitrary regions (needs a real CV dependency — see the decision record above); sub-pixel corner refinement at native resolution; a single tracked region spanning SEVERAL non-adjacent kept spans (the window is still clipped to its first kept portion, exactly as overlays are, so a plate cut into two pieces composites into the first only); an insert painted over a crossfade's blended frames (see above); lighting/color match of the insert to the scene (the content is composited as rendered — no ambient wrap, no relight); motion blur on fast plate movement.

Several of Phase 5's deferrals have since been lifted — non-planar (curved/bent) plates and multi-plate-per-frame tracking, see Non-planar surfaces and multi-plate compositing; occlusion, see Surviving occlusion. The statement below that the composite is a quadrilateral is also no longer true of the default path: it is cut to the plate's actual silhouette, see Silhouette matting. The v1 statement above that "all four corners are tracked independently" so translation, scale, rotation and perspective skew all follow the plate remains exactly true of the single-quad path, which is still what a flat plate uses.


Non-planar surfaces and multi-plate compositing​

Two limitations Tracked screen inserts (Phase 5) named as deferred, lifted. Phase 5 fits ONE planar quadrilateral per plate, per frame, which is exactly right for a flat phone screen held in frame and wrong for anything that bends — a curved monitor, a flexed card, a plate wrapped on a cylindrical object. And it takes only the LARGEST chroma component per frame, so two devices held up together resolve to one flickering track instead of two.

Both additions are off or inert by default (MaxInsertPlatesPerFrame = 1, InsertSurface = Auto with MinInsertCurvature = 0.01 that a flat plate stays well under), and a plate that measures flat keeps the original single-perspective corner-pin filtergraph.

What a chroma silhouette can and cannot tell you​

A uniform-colour plate gives exactly one thing: its outline. That bounds the problem honestly.

  • The outline does show a bend. A flat rectangle in perspective has four straight edges; the same rectangle bent about an axis has two curved ones. That curvature is measurable from pixel statistics alone.
  • The outline does not show texture foreshortening. Where a point on the interior of a curved surface projects to depends on the surface's 3D pose, and a plate of uniform colour has no landmarks to recover it from. No amount of silhouette processing changes that.

So this is an image-space deformation model, not a 3D reconstruction, and the guarantee it makes is the one the silhouette can actually support: the composited content covers exactly the region the plate covers. That is the error that shows as green fringe, and it is the one that matters. The interior distribution is smooth, temporally stable and plausible, but is not physically exact for a strongly curved or steeply yawed surface. Measured against analytically known cylindrical geometry (a rectangle bent on a circular cylinder, perspective projected):

CaseBoundary error, flat quadBoundary error, meshInterior error, flat quadInterior error, mesh
Gentle bend, distant camera3.9px0.01px3.9px0.8px
Strong bend, distant camera14.2px0.02px14.4px9.1px
Strong bend, close camera40.1px0.87px39.5px5.7px
Convex bulge (can/bottle)19.2px0.43px24.1px19.8px
Bend + 25° yaw13.1px0.51px39.6px37.4px
Bend + 45° yaw10.3px0.43px64.5px63.6px
Horizontal-axis bend18.7px0.12px18.7px0.4px
Extreme 70° bend, close camera54.0px1.29px53.0px1.5px

Read the table honestly: boundary error is fixed in every case (54px → 1.3px worst), interior error is largely fixed for fronto-parallel bends and barely improved under strong yaw. The yaw rows are the model's stated limit, not a bug to be found later.

Why not real markerless tracking​

Unchanged from Phase 5's decision record, and for the same reasons: OpenCV via OpenCvSharp has no musl/Alpine binaries the WorkflowEngine image could consume without building it from source, which means hundreds of MB of image growth and a large native attack surface parsing untrusted media, and a from-scratch pyramidal Lucas-Kanade + RANSAC feature tracker could not honestly be claimed robust here. Nothing about extending the tracker to bent surfaces changes that calculus — if anything it strengthens it, since the work below shows how far a silhouette alone can be pushed. If true 3D pose (surface normal, depth, real foreshortening) is ever required, that does need a real CV dependency, and this design does not pretend otherwise. What it buys instead is the capability that a bent plate composites cleanly, with the constraint stated: the footage must contain a uniform-colour plate, and the interior mapping is an approximation.

Detection: ChromaPlateEdgeProfiler​

Per frame, per plate, against the mask the tracker already built. For each of the four edges it walks the edge in the quad's own projective parameter t — the same parameterization InsertSurfaceMesh replays, so measurement and reconstruction cannot disagree — and at each sample marches from a seed known to be inside the plate outward until the mask stops being set. That gives a signed deviation profile in units of the edge's own "across the plate" vector: dimensionless, so a number measured on the 320x180 tracking grid replays at 1080p with no conversion.

The measurement is two passes, and real footage is why. The first version pinned the deviation profile to zero at both ends, deliberately refusing to move the tracker's corners so that curvature measurement stayed separable from corner-fitting accuracy. Run against the project's real clips that was simply wrong. Every real screen has ROUNDED corners, and the extreme-point corner fit lands on the rounding's 45° tangent point, several pixels inside where the straight edges meet; measured against those inset corners, a perfectly flat plate looks bowed on all four edges. The real phone clip read 0.027 of curvature against a 0.01 threshold, and a barrel-distorted version of the same clip read 0.025 — flat and curved indistinguishable, with the flat case over the line.

So: pass one fits a FREE quadratic to each edge over its middle band only (15% dropped at each end, exactly where rounding lives) and extrapolates back to the endpoints; each corner moves by the sum of the two contributions from the two edges meeting there, which are roughly orthogonal and so compose into a proper 2D correction. Pass two re-probes against the refined quad with the profile pinned at the ends again, and what survives is the genuine bend. A rounded-corner rectangle comes out "corners nudged outward, edges straight"; a curved plate comes out "corners roughly unchanged, edges bowed". The real clips now read 0.007 (flat) against 0.056 (barrel-curved).

The reading is grid-resolution dependent, so cross-check a near-threshold one​

Measured while choosing footage for the alpha launch film, on two real unmodified clips (a phone held in frame, and a desktop monitor showing a full green screen with hands typing in front of it), by running the same analyze step twice and changing only InsertGridWidth/InsertGridHeight:

ClipAuto grid (0/0, ~320px long axis)Forced 480x270
desktop monitor0.0065 → Flat0.0179 → Curved
phone0.0056 → Flat0.0133 → Curved

Both plates are physically flat screens, and both cross the 0.01 threshold purely by being probed on a denser grid. The mechanism is the obvious one — a finer mask resolves more of the anti-aliased edge, and the residual bow the refined-corner pass cannot absorb grows with it — but the practical consequence is worth stating: a curvature reading just above MinInsertCurvature is not evidence of physical bend. Re-probe it at a second grid resolution; a genuinely curved plate reads curved at both, while a flat one flips. Auto is also simply the right default for a mixed portrait/landscape source set, since it preserves each source's own aspect per the tracking grid section.

The bow model is two parameters per edge — dev(t) = 4t(1-t)(bow + skew(2t-1)), a symmetric bulge plus the leaning one a yawed bend produces. Deliberately not a free-form boundary: a free-form fit would follow a phone's notch, a thumb crossing the bezel, or a compression artifact, and the Coons blend would ripple that dent through the whole interior. The two-parameter family is structurally the wrong shape to express a notch, so notches fall out as fit outliers instead of becoming geometry — and it is four numbers per frame, which keeps the track temporally stable and the artifact small.

Two validity gates, both forced by real footage​

The project's laptop clip defeats the extreme-point corner fit outright: the plate is a strong trapezoid, "top-left" (min of x+y) lands partway down the left edge, and the fitted quad cuts the entire upper-left triangle off the plate — at confidence 1.000 and fill ratio 1.000, a bad-but-confident detection. The silhouette then genuinely does sit far outside that quad's top edge, and the profiler dutifully measured a 0.118 bow. Meshing on that reading made the composite visibly worse than the flat pin: it bent the top edge into a parabola that matches a triangle nowhere and crushed the interior. Two independent gates now catch it:

GateRejects whenFlat phoneBarrel-curvedLaptop (broken fit)
Corner validitypass one wants to move a corner more than 0.10 of its across-vector — further than rounding could ever explain0.0250.0550.141
Fit residualthe pinned profile's RMS residual exceeds 0.012, i.e. the model does not describe this boundary0.0050.0040.036

The residual tolerance is deliberately ABSOLUTE and not scaled by the bow's size: the residual floor is grid quantization, which does not grow with curvature — the barrel clip carries four times the bow at a lower residual than the flat one. Either gate firing discards the plate's curvature and keeps the tracker's own corners, so the insert falls back to the corner pin, which is no worse than before. The laptop clip now reads 0.000 and composites exactly as it did.

(The underlying corner-fit failure is a Phase 5 tracker bug, not one this addition introduces — noted here because it is what the gates exist to survive.)

Geometry: InsertSurfaceMesh​

The interior is the projective map H of the four (refined) corners — identical to what the single-quad perspective path already applies — plus a transfinite (Coons) blend of the four measured boundary deviations:

P(u,v) = H(u,v)
+ (1-v)·devTop(u) · (H(u,1) - H(u,0))
+ v ·devBottom(u) · (H(u,0) - H(u,1))
+ (1-u)·devLeft(v) · (H(1,v) - H(0,v))
+ u ·devRight(v) · (H(0,v) - H(1,v))

Two properties make this the right base. All-flat bows reproduce the corner homography exactly, at any mesh density: a projective map is uniquely determined by four point correspondences, so a cell whose corners are sampled from H refits to H itself. A mesh over a flat plate is therefore the same warp as one big perspective, which is what lets the executor route flat plates to the original builder rather than approximating it. And every deviation vanishes at the corners, so the Coons construction's four corner terms are identically zero and the blend needs no corner-correction term.

Density is derived, not configured: a parabolic bow of peak s split into n cells leaves a per-cell chord error of s/n², so n = ceil(sqrt(s / InsertMeshTolerancePx)) is the exact subdivision that brings the straight cell edges within tolerance of the measured curve. Top/bottom bows drive columns, left/right bows drive rows, and MaxInsertMeshCells caps the product. The density is chosen once per insert from the MEDIAN bow across its keyframes, so one badly-read frame cannot inflate the graph and each cell's canvas can stay fixed for the insert's duration.

Rendering: ScreenInsertMeshFilterBuilder​

A single projective transform cannot represent a non-planar surface at all, so the surface is tiled into Rows x Cols cells, each its own perspective warp.

remap was the alternative and was rejected on measured grounds. ffmpeg's remap takes an arbitrary per-pixel coordinate map and would need no tiling whatsoever — but it samples NEAREST-NEIGHBOUR from integer gray16 maps, which shimmers badly on exactly the content this feature exists to composite (a UI with text, scaled down into a phone-sized region), and it would require generating and decoding gigabytes of raw per-frame map streams: a new I/O mechanism, where the tiled form stays what every other filter builder in this feature already is, a pure string. The tiling costs encode time instead, and that was measured as cheap — a 3x6 mesh at 1080p ran at 94 fps against 253 fps for the single-quad path.

Per cell, on its own small canvas sized to that cell's destination bounding box over the whole insert (not the full frame — that is what keeps a dense mesh affordable and keeps each source cell resampled at roughly its destination scale):

[asset] fps, format, setpts(+outputStart), split into Rows*Cols
[cell] crop to this cell's source sub-rectangle (floor-tiled, so cells abut exactly),
scale to (cw-2)x(ch-2), pad to cw x ch with a 1px black border,
perspective sense=destination eval=frame -> warped cell
color=white (cw-2)x(ch-2), pad 1px black, format=gray, SAME perspective -> warped cell mask
alphamerge ; overlay at the cell canvas origin, enable='between(t,start,end)'

Seams. Two mechanisms, and the second was found by rendering real footage rather than reasoned out. First, perspective is handed not the cell's mesh corners but the exact projective EXTRAPOLATION of them out to the padded canvas corners — the content sits 1px inside the canvas because of the border the warp needs, so asking it to send the CANVAS corners to the mesh points would land the content a border-pixel short on every side and open a gap at every interior boundary. Second, real renders still showed a green hairline along every interior boundary: perspective interpolates bilinearly, so both the warped content and the warped white mask fade out over roughly a pixel at the cell edge, and two abutting cells each contributing ~50% alpha composite to ~75% coverage with the plate showing through the rest. Each cell is now grown 2px onto every side that HAS a neighbour, so one cell's fully opaque interior covers its neighbour's ramp. The mesh's outer boundary is never grown — it must stay exactly on the measured silhouette (InsertOverscan remains the separate, deliberate knob for spilling past the plate edge).

Mesh and flat inserts interleave freely inside one chain: ScreenInsertFilterBuilder dispatches per insert and the mesh builder ends at the same label the flat branch would have, so the compile step's input bookkeeping, stage ordering and label plumbing are untouched. A mesh insert is still ONE extra ffmpeg input however many cells it has.

Multi-plate ("multipoint") tracking​

The interpretation pursued here is several independent chroma plates visible at the same time — two devices in one shot, or a device and a separate monitor — each becoming its own r{n} region with its own rendered asset. A single physical object with two rigidly-hinged faces (an open laptop's screen and keyboard deck) falls out of the same mechanism as two plates; what is NOT modelled is any constraint linking them, so their tracks are independent and could in principle drift relative to each other.

ChromaQuadTracker now keeps the MaxPlatesPerFrame largest components clearing the area/fill gates, in descending area order with ties broken by raster position, and associates them across frames greedily by centroid distance — all candidate (track, detection) pairs in ascending distance, ties broken by index, so the assembly is bit-for-bit reproducible. Tracks unmatched for longer than MaxGapFrames close; unmatched detections open new ones. At the default cap of 1 there is at most one detection per frame and this reduces to the original single-run walk, with one deliberate improvement: a track stays open across the gap window, so a plate that jumps away and returns within it is re-associated instead of split.

The compile side needed no changes at all — ResolveInsertsAsync already looped over plan.Inserts and chained one composite per region, so simultaneous regions "just work" once the analyze step produces them.

Agent surface​

Unchanged in shape. view.insertRegions gains one descriptive word, surface: "Flat" | "Curved", so the planner can render content suited to a bent screen. The bow numbers, the refined corners and the per-frame tracking data stay compile-time-only in the artifact — the same discipline as every other id family. ScreenInsert is untouched: still exactly three strings, still pinned by MotionGraphicsPlanOutputInvariantTests. A curved surface is a property of the SHOT, measured by first-party C#; it is not something an agent may assert, request or influence.

Config reference​

VideoAnalyzeStepConfigDefaultMeaning
MaxInsertPlatesPerFrame1Chroma plates tracked simultaneously per frame. 1 is the original behaviour exactly. Clamped 1..16; MaxInsertRegions still caps the artifact total
MeasureInsertCurvaturetrueMeasure refined corners + edge curvature. false reproduces the pre-existing composite exactly
VideoCompileStepConfigDefaultMeaning
InsertSurfaceAutoAuto meshes a plate whose measured curvature clears the threshold; Planar never meshes; Mesh always tries (debugging only — on a flat plate it is the same warp, more expensively)
MinInsertCurvature0.01Curvature a plate must exceed under Auto to earn a mesh, as a fraction of its own size
InsertMeshTolerancePx0.6Target worst-case chord error per cell; this is what sets the cell count
MaxInsertMeshCells24Ceiling on rows * cols for one insert

Real-footage results​

Validated against the project's own green-screen clips (project 2f9216f9-dc7d-4118-9e60-e2855cb38218), driving the shipped ChromaQuadTracker → VideoCompileStepExecutor → ScreenInsertFilterBuilder path, composited with a solid-colour asset so "is this pixel content?" is exact. Registration is scored TWO-SIDED against the source frame's own chroma mask, because green-leak alone rewards simply over-covering — which is precisely how a flat pin can hide curvature:

  • leak — plate pixel NOT covered by content: green showing through, the visible failure.
  • spill — content pixel NOT on the plate: the insert painted onto the bezel or the hand.

Absolute spill is large for every variant and is not by itself error: InsertOverscan deliberately expands the quad by 2%, and the source mask is conservative at the plate's anti-aliased edge. Only the DIFFERENCES between variants mean anything.

ClipVariantleakspill
greenscreen-phone.mp4 — flat plate, rounded corners + notch (reads 0.006, stays planar)fitted corners0.55%39.7%
refined corners0.22%43.8%
the same clip through a known barrel distortion (reads 0.033 → mesh 5x4)fitted corners, flat pin0.24%41.1%
refined corners, flat pin0.16%46.5%
refined corners, mesh warp0.18%41.3%
greenscreen-laptop.mp4 — flat plate, strong trapezoid (reads 0.015)either path~0.00%—
both real clips composited into one frametwo concurrent regionsboth plates composited independently, each tracking its own motion

Read plainly:

  • Corner refinement is the clear win, and it is a trade. It more than halves leak on a real flat plate (0.55% → 0.22%) by pushing the quad out to where the straight edges actually meet, and pays for it in a little more content on the bezel. That is the right side of the trade — green showing through a screen is glaring, a few pixels over a dark bezel is not.
  • On a curved plate the mesh buys back exactly that cost. At the same leak it drops spill from 46.5% to 41.3%, i.e. it CONFORMS to the plate where the flat pin over-covers it. Refined corners plus mesh is strictly better than the original on both axes (leak 0.24% → 0.18%, spill unchanged).
  • The mesh does NOT reduce green on this footage, and it would be wrong to claim it does. With refined corners and 2% overscan the flat pin already covers a mildly curved plate; what the mesh improves is registration, not coverage. Its coverage advantage is demonstrated against analytically known geometry (the table further up: boundary error 54px → 1.3px), not against any clip in this project.

The barrel-distortion clip is the honest part of this to be clear about: it is real footage put through a known geometric distortion, which is what a wide-angle lens does to a flat plate, not a camera-captured curved object. It exercises the capability on real pixels with real motion, compression and noise, and a genuinely curved silhouette. Footage that does not exist in the project, and that would be needed to validate what is left: a physically curved screen or a plate on a cylindrical object — with enough bend that a flat pin cannot hide it under overscan — and a single shot containing two devices at once (the two-plate result above is a spatial composite of two real single-plate clips).

Sourcing a genuinely curved plate: searched, not found​

Looked for while producing the alpha launch film, across Pexels, Pixabay, Mixkit, Videezy, Hollywood Camera Work's free VFX plates, and an ID sweep of one studio's fabric series, under every phrasing that seemed likely: green/blue/magenta cloth, fabric, flag, banner, curtain, tarpaulin, veil, drape, blanket, morphsuit, cylindrical product wraps, held cards and posters.

The finding is consistent enough to be worth writing down. In free stock libraries a chroma colour appears in exactly two forms: as a backdrop behind a subject, or as a full-frame fabric texture shot edge to edge. Both are useless here for the same reason — the tracker needs the plate's OUTLINE, and neither has one. The near misses illustrate it: studio fabric series shoot beautiful billowing cloth as an outlined foreground object against black, but in white, red, yellow and cream; the one magenta clip in that series has the magenta as the BACKDROP with a white cloth in front of it. Green morphsuit performers exist against black, but a head-plus-torso-plus-arms silhouette is not a quadrilateral and the validity gates reject it, correctly.

So a camera-captured curved chroma plate was NOT obtainable, and synthesising one — warping a flat clip to manufacture a passing curvature reading — would be fabricating the test data the gates exist to evaluate. If this capability is to be validated against real footage, the realistic route is to SHOOT it: a green cloth taped over a cylinder, or a sheet of green card bowed between two hands, is a few minutes of work with any phone and would settle the question properly.

Limits, stated​

  • Interior foreshortening is approximate under strong yaw or strong convexity — see the table above. The silhouette does not carry the information; only real 3D pose estimation would.
  • A convex plate's corners are biased. When a plate bulges outward, the extreme-point corner fit lands on the bulge rather than at the true corner, and the bow is measured relative to that. The corner refinement absorbs part of it; the validity gate rejects the rest rather than guessing.
  • Corner-fit failures are survived, not repaired. The laptop case degrades to the Phase 5 behaviour. Fixing the extreme-point fit itself is separate work.
  • No inter-plate constraint. Two plates on one rigid object track independently.
  • Everything Phase 5 deferred that is not listed above stays deferred: occlusion recovery beyond gap bridging, sub-pixel corner refinement at native resolution, inserts on multi-source or crossfade compiles, lighting/colour match of the insert to the scene, motion blur.

Silhouette matting​

A tracked insert has to be cut to the shape of the plate it goes into, not to the quadrilateral that approximates it. Phase 5 and the non-planar work above both answer where the content goes; neither answers which pixels it covers, and until this section's work every composite was a filled quad. On the project's own footage that meant the phone's rounded bezel corners and its camera notch were painted over, and a hand crossing the laptop's screen was painted over too.

A real plate's silhouette is not a quadrilateral and cannot be made into one by bending its edges. It has to be cut out, per frame.

warp decides WHERE the content lands (quad / mesh, from tracked corners)
matte decides WHICH PIXELS it covers (this section)
spill suppression decides WHAT COLOUR survives (this section)

The cutout already exists in the source​

The plate is a known colour, so keying the base picture gives the matte directly: pixel-exact, at full encode resolution rather than the coarse tracking grid, anti-aliased by the source's own sampling, and free of any extra decoding. Rounded corners, notches, an occluding hand and the motion blur on that hand all come along with it, because they are all simply "not the plate colour".

The insert's alpha therefore becomes a conjunction:

alpha = warped quad/mesh mask AND plate matte keyed from the base

Both halves are load-bearing. The quad mask keeps green scenery elsewhere in frame out; the plate matte keeps the bezel, the notch and anything in front of the plate in.

Why the key is searched rather than configured​

A key needs a colour and a tolerance, and no fixed pair works. The project's two clips measure the plate at RGB (0, 232, 0) — bright and saturated — and (31, 63, 10) — dark and nearly neutral. A tolerance that keys the second keys half a room in the first. ffmpeg's chromakey, which ignores luma and compares chroma only, fails outright on the dark plate for exactly that reason: its chroma sits close to the rest of a dim room.

Rather than expose a knob nobody can set without looking at frames, PlateChromaKeySolver grid-searches colour × similarity × blend per plate and scores each candidate against ground truth it already has: the plate mask ChromaQuadTracker segmented for that very frame.

ElementHow it is chosen
Candidate coloursPercentiles of the plate's own luma distribution (25/40/50/60/75) — percentiles, not a mean, because a plate routinely carries a specular band and a dark fringe whose mean is a colour appearing nowhere on it
SimilaritySwept across the full usable range, finely enough that the winner is not a coarse compromise
BlendSwept as fractions of the chosen similarity, so the edge ramp scales with the key's own tolerance
ScoreSoft IoU (Σmin/Σmax) against the tracker's segmentation, plus a spill term (below)

The distance metric replicates ffmpeg's vf_colorkey exactly — diff = sqrt((dr²+dg²+db²)/3)/255, alpha 0 at diff ≤ similarity ramping to 1 by similarity + blend — verified against a real ffmpeg run over a synthesized distance ramp, because a score computed against a different curve than the one that will run would be worthless.

Candidates are judged only inside the plate's dilated bounding box. That is the only place the matte decides anything: the warped quad mask already excludes everything else, so a candidate that keys foliage across the room is not penalised for it, while one that eats into the bezel — the failure that matters — is.

Measured on the real clips, with no clip-specific constant anywhere:

ClipSolved keySimilarityAgreement
greenscreen-phone.mp4#00EC000.3680.77
greenscreen-laptop.mp4#193E080.0560.98

The spill term​

Where something crosses in front of a plate, the pixels along its edge are a physical blend of that object and the plate behind it. They carry real green, more of it the more motion blur widens the edge. A key tight enough to match the tracker's segmentation leaves them in the base, and the composite then shows a coloured rim tracing the occluder.

Plain agreement cannot fix this: that rim is a handful of pixels against a whole plate and rounds away in the IoU. The objective therefore carries a second term that measures the miss normalised over the blend band itself, giving those pixels weight proportional to the problem they cause. The band is the difference between the tracker's own two detection thresholds — green enough for the permissive grow test, not for the strict core test — restricted to inside the tracked quad so that green scenery behind the plate can never drive the key wider.

Two deliberate details:

  • Selection and reporting are different numbers. The spill term steers the choice; the stored Score stays plain agreement. Folded into one number, a plate with an inherently wide fringe would have its matte switched off for having a hard problem rather than for being badly solved.
  • The floor is lexicographic. Any candidate clearing MinUsableScore outranks any candidate that does not, whatever their spill — otherwise a heavy spill penalty could hand the win to a key whose agreement had fallen under the floor, which reads downstream as "no matte at all".

A plate whose best candidate still disagrees with the tracker stores no key, and the compile step then composites with the plain quad mask exactly as it did before matting existed.

Spill suppression​

Matting cannot fix a pixel it is correctly keeping. Once the key is as wide as it can go without swallowing the occluder, what remains has to be corrected photometrically:

G' = min(G, max(R, B)) the classic despill
= max(min(G,R), min(G,B)) the identity actually built, from two channel-mixed copies
and a darken/lighten pair — no plane surgery, R and B untouched

It is, by construction, a no-op on any pixel where green is not dominant, so it cannot discolour skin, a bezel or a wall. Verified on real pixels: plate (28,66,10) → (28,28,10), skin (145,103,69) → unchanged.

It must run in planar RGB, and for a long time it did not. blend has no packed-RGB support, so ffmpeg converts its inputs to something it does support, and inside a full compile graph that negotiation picked a YUV format, where darken and lighten are no longer per-RGB-channel min and max. Measured on the phone clip: (33,73,32) came out as (6,50,5) rather than (33,33,32), so suppression barely suppressed, and the rim it exists to clean stayed green. Run on its own, the same chain happened to negotiate gbrp, which is why it looked right in isolation. Every branch of it is now pinned with format=gbrp. Residual green round the phone insert fell from 61-113 px a frame to zero on the colorkey matte path.

The corrected picture is laid over the base with the band as its alpha, not maskedmerged. The band is synthesized a margin past the insert's window, and maskedmerge repeats every input's last frame until the longest one ends (this ffmpeg has no shortest option for it). The finished film therefore ran about a second past its source on its held final frame, after the insert's enable window had closed, and any video whose insert reached its last frame ended on a raw green screen. An overlay always ends with its main input, the base picture.

It is still bounded, because the one thing it would wrongly touch is genuinely green scenery. The bound is the insert's own warped footprint minus the matte — precisely the occluder where it crosses the plate, at whatever width its motion blur happens to be. Two other bounds were tried and are worth recording as dead ends: a fixed dilation radius is too narrow for a fast-moving edge and too wide for a locked-off one, and deriving that radius from the plate's blend statistics fails because the region those statistics describe (grow-minus-core) is the plate's dim interior, not its edge — it measured 145 pixels on a clip whose edge is one or two.

Overscan is inverted by matting​

Without a matte, expanding the quad past the plate paints content onto the bezel, so InsertOverscan defaults to a cautious 2%. That leaves the content stopping a hair short of the plate wherever the coarse tracking grid rounded a corner inward — visible on real footage as a green seam down the plate's left and bottom edges.

With a matte the content is clipped to the plate's silhouette however far it is expanded, so overscan becomes free and generous is strictly better: it guarantees the content reaches every pixel the matte will admit. A matted insert therefore uses at least 10%, never reducing a larger configured value. Measured: residual plate visible inside the composite fell from 2066 px to 545 px on the laptop occlusion frame.

How the content reaches that band changed later. It used to be padded by the same factor (black) and warped onto the expanded quad, which lands the content's own frame back on the plate only for an unforeshortened plate; a flat matted insert now pins the content's frame to the tracked plate itself and lets the warp smear its edge pixels across the band — see "Smeared overscan" in Tracking quality on the customer clips.


Sub-pixel corners at the source's frame rate​

A tracked insert wobbled, visibly, even when its outline followed the plate well. On the project's generated phone clip (1344x768, 24 fps), measured against a full-resolution reference track built from edge-line fits, the corners the compile step consumed sat 3.0 px off the true corners on average and moved 3.7 px RMS (6.8 px at the 95th percentile) around that offset, every corner independently, which is what makes inserted content shear. Three separate causes, each fixed separately:

CauseFixWobble after
Each corner was ONE grid pixel (the extreme of x±y), quantized to ~4 source px and landing on the corner rounding, ~5 px inside the true cornerChromaEdgeLineFitter: each edge located as a line through 40 sub-pixel crossings, corners where the lines meet1.9 px
The grid was sampled at 10 fps from a 24 fps source, so every sample was the NEAREST source frame, up to 21 ms off its stamped timeSample at the source's own rate, or an exact divisor of it (InsertSampleFps = 0, auto)1.6 px
The default grid was 320 px wide480 px (ChromaQuadTracker.AutoGridBaseWidth)1.4 px

The reference itself is noisy at about 1 px, so the result is close to what can be measured. The offset from the true corners fell from 3.0 to 1.0 px.

The edge fit. Along the middle 80% of each edge (corner rounding lives at the ends), a probe marches across the edge through the bilinearly interpolated dominance signal, target channel minus the strongest rival in absolute 0..255 units, and records where it falls through half the plate's own level. That 50% point of the soft plate/bezel transition is located from the intensity ramp to a fraction of a grid pixel. Absolute dominance rather than the brightness-normalised purity the growth bar uses, because the plate is usually framed by a near-black bezel, where purity is a ratio of small numbers and is noise: fitting to purity put every edge 3-5 px out on the bezel. A Tukey-weighted total-least-squares line through each edge's crossings ignores the stretch a reflection, a fingertip or a notch displaced, and a probe that does not start inside the plate is skipped rather than guessed. The fit returns nothing, and the extreme-point corners are kept, when an edge has too few crossings, a corner would move further than 12% of the diagonal, or the result is not convex.

Flat plates skip the profiler. The same fit measures how straight each edge is (the bulge of a quadratic through its residuals, as a fraction of the plate's extent across it). A plate whose edges are all straight (FlatEdgeBowFraction, 0.025) is fully described by its four line-fitted corners, and those become its surface corners too. Previously ChromaPlateEdgeProfiler re-refined every plate from the binary component mask, read grid ripple on the flat phone as a bow on most frames, and moved the corners the compile step uses, adding its own ~1.5 px of wobble. Measured on the phone: median 0.007, 90th percentile 0.011. A genuinely curved plate (the synthetic barrel test reads 0.14) still goes through the profiler and keeps its mesh warp.

Frame-count knobs keep their meaning in seconds. MaxGapFrames, MaxOccludedFrames and the occlusion and consistency windows were tuned at 10 fps; Track rescales them by fps / ReferenceSampleFps, never below 1x. The 3-frame smoothing window is deliberately left in frames: it averages per-frame measurement noise, and at a higher rate it lags real motion less.

Memory. The insert grid is read whole (the tuner and assembly both walk it), so the automatic rate is chosen inside a fixed 640 MB budget (InsertGridByteBudget), about what the old 3000-frame cap cost at 320x180. A 6.6 s clip samples every frame at 480x274. A long clip drops to the largest exact divisor that fits (a two-minute 24 fps clip samples at 12), and only when even 2 fps would not fit does it fall back to an arbitrary rate.

Dense keyframes reach the encode. MaxInsertExprKeyframes used to thin every track to 96 keyframes, roughly 4 fps on a 24 s insert, reintroducing the interpolation error dense sampling removes. Past 96 keyframes the corner series is now written as a balanced tree of if(lt(in,…)) splits (log-depth nesting, which ffmpeg's expression parser accepts at any length), and the default cap is 1500. Up to 96 keyframes the filtergraph is byte-identical to before.

A weak fragment at a track's end is dropped (ChromaQuadTracker.TrimWeakEnds). On the phone clip the phone is turned toward the camera through a flash of glare: two frames see a thin sliver of the screen edge-on, the next five show a pale reflection with no green at all, then the plate is clean. Gap bridging joined the sliver to the body, so the insert opened on the sliver and interpolated across the glare while the phone rotated: a quad sliding off the phone for a quarter of a second. A markerless feature tracker (the Tracking/ LK + RANSAC pipeline) was tried for those frames and loses the phone too, to fast rotation, motion blur and a screen whose appearance flips from green to white. With no trustworthy geometry, the right composite is none. A leading or trailing run of detections that is separated from the track's body by a bridged gap, and whose median core area is under half the body's, is removed together with the bridge. A plate that starts small and grows without a gap keeps every frame.

Tracking through a turn​

After the sub-pixel work above the insert sat still on a still phone, and still wobbled while the phone was being TURNED toward the camera — the 20 frames of the project's phone clip where the plate is foreshortened, motion-blurred and crossed by reflection streaks. Three things were wrong, none of them "noise", and the first job was a measurement that could see them.

A reference for the turn. The earlier reference (edge-line fits to the green at full resolution) broke on exactly the frames that mattered, because the green does. The reference used here is bounded by the phone's dark BEZEL instead: from a temporal prior, 60 probes per edge march outward through the full-resolution frame and take the dominance crossing where the probe starts on plate colour, or the luminance drop into the bezel where it starts on a white reflection — the two signals fail on different pixels, so together they cover the glare frames the green-only reference could not ($S/gt_gen.py, propagated from the frame whose detection best agrees with its neighbours). Its per-edge residual is 0.3 px and its second differences on the steady part imply ~0.4 px of noise. The metric reports, per corner, the error against it and the slip — the frame-to-frame change of that error, which is content sliding on the glass and is what the eye reads as wobble. Both were checked against overlays of the quads on the frames before anything was optimised against them.

Diagnosis. Against that reference the per-frame detections were within ~1 px on EVERY frame of the turn except three: two where a reflection streak crossed the top and left edges and the edge-line fit followed the streak (63 and 41 px off), and the first frame of the track, where half the screen was glare and the fit bounded only the green half. The 3-frame average then spread each of those over its neighbours — a frame whose own detection was 0.9 px off came out 27 px off — which is why the insert looked wrong for six frames when three were measured wrong. Smoothing was the wrong tool; a corrupted measurement has to be recognised and replaced.

ChangeWhyWhere
Corners in the plate's own frameThe extreme-point corner rule (TL = min x+y, …) assumes an upright plate; past ~30° two extremes converge on one physical corner and at 45° the "quad" is a triangle. Generated test shots turn a tablet from landscape to portrait and whip a phone up through 45°, and every frame of those turns fitted a kite. The orientation of the minimum-area rectangle around the component's convex hull is measured first and the rule applied to coordinates rotated by minus that angle; upright plates are unchangedChromaQuadOrientation
Labels follow the physical cornersThat orientation is only defined modulo a quarter turn, so the fit renames the corners as a plate passes 45°, and content pinned to "top-left" would snap a quarter turn. Each matched detection is re-labelled by the cyclic shift closest to the track's previous sample (surface corners and edge bows shift with it). Continuity alone inherits whatever labelling the FIRST detection got, a coin toss when that frame sits near 45° — on the whip clip it put the content sideways on an upright phone for four seconds — so one further shift is applied to the whole track afterwards, the one that makes the "top" edge the most nearly horizontal over the track's median frame: the pose the plate holds longest is the one that reads upright, and a plate turning landscape → portrait gets whichever it held longer, continuously through the turnChromaQuadTracker.AlignCornerLabels, UprightTrackLabels
Implausible detections are replacedEvery detected sample is predicted by a straight line through its detected neighbours within ±125 ms; the furthest any corner sits from that prediction is compared with the larger of 2% of the plate's own diagonal and 1.5× the plate's own local per-frame travel. Both arms are relative to the footage — the floor scales with the plate, the speed arm with its motion — so a tablet filling the frame and a phone in the distance, a whip and a locked-off shot, are each judged against themselves, and a generated clip's frame-hold-then-double-step (one frame's travel off a straight line) passes. Rejection is greedy, worst first, re-predicting after each removal: judging every frame against predictions that still contain the outlier flagged all of the turn and then none of it. Rejected interior frames are interpolated from their clean neighbours; a rejected END frame is dropped. Capped at 30% of a track's detectionsChromaQuadTracker.RejectImplausibleSamples, Options.OutlierFloorFraction / OutlierSpeedFactor / MaxOutlierFraction
Smoothing is weighted by motionThe 3-frame average removes noise of the order of the measurement noise and nothing else; on a plate moving several pixels a frame it lags the motion by a fraction of a frame's travel and smears generated footage's frame holds. Each frame is blended between its own detection and the window's average with weight 1 / (1 + (motion / (2 × noise))²), motion being its neighbours' median per-frame travel and noise the track's own median residual from the rejection pass: a still plate is fully averaged, a plate moving at twice its noise half, a fast one passed throughChromaQuadTracker.Smooth, Options.SmoothingMotionGate

Measured (corner error / slip against the bezel reference, source pixels; $S/metric_gen.py). Before is the tracker as shipped above; after is all four changes:

Clip (generated with MiniMax, 24 fps)FramesBefore: err mean / worst / slip RMSAfter
Phone turned toward camera (1344x768), the turn, frames 26-45204.46 / 35.2 / 7.011.48 / 9.5 / 2.72
Same clip, steady part1121.20 / 5.0 / 1.231.18 / 2.9 / 1.01
Blue-screen phone tilted outdoors against foliage, handheld1581.13 / 4.0 / 1.680.86 / 3.6 / 0.62
Laptop on a desk, slow dolly, dim plate1580.80 / 2.9 / 0.310.81 / 2.9 / 0.35
Phone whipped up through 45° to fill the frame (768x1344) — frames where the plate is in picture981.09 / 12.1 / 0.760.99 / 3.7 / 1.04
Tablet turned landscape → portrait by a window15623.8 / 409 / 39.419.0 / 340 / 6.5

Variants tried on all six clips ($S/sweep_table.txt), each change isolated: the orientation fix alone changes only the two turning clips (tablet slip 39 → 7); rejection alone takes the phone turn from 4.46 to 2.46 px but leaves the average's lag (slip 4.6); the gate alone does nothing measurable on the phone and is what fixes the blue clip; rejection without any smoothing is as good on the turn (1.39) and worse everywhere still (phone steady 1.31 / 1.27, whip slip 1.74). A hard on/off gate at 2× noise instead of the weight above was equal or worse on every clip (whip slip 1.30 vs 1.04); 1× and 4× were each worse on at least two clips; a 5-frame window was worse on the whip and the blue clip. Stricter rejection (1% / 1.0×) rejects frame holds and costs the whip its worst frame back (11.9 px); laxer (4% / 2.5×) is identical on the phone and worse on the tablet's slip. A finger-swipe clip whose plate runs off the top and bottom of the picture is not in the table: the reference itself fails where the finger covers an edge, and that clip is judged by eye (its right edge still follows the finger — occlusion of an EDGE, as opposed to the area collapse "Surviving occlusion" detects, was an open problem here; it is fixed in Tracking quality on the customer clips).

The two kite cases are where the orientation fix shows: the whip's worst frame goes from 12 to 4 px, and the tablet, which no longer fits a kite between 30° and 60°, drops its slip six-fold. The tablet's remaining error is a different problem, left open: the window reflection on its glass is green-TINTED, not white, so it sits below the purity bar the component grows through, the component excludes it, and the coarse right edge — and the line fit that starts from it — lands on the reflection's boundary inside the screen. The laptop is a plate that never moves, where the gate is never shut and the numbers are unchanged within the reference's own noise.

What did not help. Replacing the moving average with a local quadratic (Savitzky–Golay, 5 or 7 frames) was worse on the turn (2.2 px) than either the gated average or no smoothing at all, because the real motion is a staircase of held and doubled frames that no polynomial follows. Smoothing in homography space and a Kalman smoother were not tried once the diagnosis showed the error was three gross outliers and not noise: no filter of the measurements fixes a measurement that is of the wrong thing. Bridging a rejected frame with a quadratic through its ±3 clean neighbours instead of a line gained 0.08 px mean and 2 px on the worst frame of the turn, but a ±4-frame window and a least-squares line through the same neighbours were both WORSE than plain interpolation between the two nearest — too fragile to ship for the gain, so a rejected frame is bridged linearly like a detection dropout. The frame-hold cadence of generated footage is itself the floor on what interpolation can do: the plate stands still for a frame and then moves twice as far, and the 9.5 px worst frame after the change is a bridged frame on that cadence.

Reflections​

A screen insert is only convincing if the glass still looks like glass. On the project's phone clip, the phone catches a reflection as it turns. Before this, the composite showed two artefacts. The reflection streak was left un-keyed, so the raw plate showed through the insert as pale green blobs. And a potted plant behind the subject was keyed as plate, so content was painted onto the plant wherever the insert's overscan reached it. Both came from the colorkey matte, and the fix has two parts: a matte keyed on the plate's own chroma level, and carrying the reflection over.

The dominance matte​

colorkey keys by RGB distance from one colour, and that radius is the wrong shape on both sides. A reflection is the plate plus white light, (150,215,160) against a plate of (0,202,0), far outside any radius that excludes the room. Dull scenery green, (70,93,26), is close in RGB distance to a dark plate. The tracker now records each plate's own median dominance (VideoInsertChromaKey.PlateDominance, 0..255), and when it is present the compile step keys on that instead. The matte is the maximum of three terms, each a function only of the pixel's dominance x and its target-channel level y, so the whole matte is ONE lut2 table lookup:

TermRamps overCatches
Platedominance 35%..65% of the plate's levelthe plate; the ramp is centred on the same 50% point the edge fit puts the edge at
Dark platepurity x/y 0.5..0.75, above an absolute floor of 10..24the shadowed plate and its 1-2 px anti-aliased edge into the bezel, (0,25,0): purity 1.0, a tenth of the plate's dominance. Without it that rim showed as a green outline
Reflectiondominance 8..24 AND target level 45%..65% of the plate'sglare on the plate: light added keeps the target channel near the plate's own (dimmest streak pixels ~121 against a plate of 202), dull scenery sits far below (the plant: 93)

White walls and the pale blue shirt have no dominance at all and never key. The plant reads about 5% at the foot of the reflection ramp, the price of keeping the dimmest real streak pixels. An artifact written before PlateDominance existed keeps its colorkey matte unchanged.

Carrying the reflection over​

With InsertReflections on (the default; it needs a matte), the plate's reflection layer is screen-blended over the composited insert, inside the matte and the insert's footprint only. Glass shows what is behind it plus what it reflects, and screen is the clipped form of that addition. The reflection layer is the ACHROMATIC part of the base picture, min(R,G,B) in every channel less a noise floor of 8. A clean plate of any hue contributes black, a no-op under screen, and white glare carries over intact. The plate's colour cannot simply be despilled away instead, because real plates are not pure: the phone's reads (0,171,30) where the glass turns from the light, despilling keeps that blue as "reflection", and over the insert's black content it tinted the whole screen dark teal. The cost of taking only the neutral part is that a coloured reflection, say a red shirt mirrored in the glass, carries over as grey glare.

Two ffmpeg details are load-bearing. The blend runs in planar RGB (gbrp): screen applied to YUV planes blends the chroma planes too and shifts every colour of the insert (red read as pink). The result is merged back with alphamerge + overlay, not maskedmerge: the screened copy differs from the composite over the WHOLE frame, and maskedmerge, given it with this mask, returned the screened copy everywhere, tinting the finished frame magenta edge to edge.


Surviving occlusion​

Occlusion is missing data, not a smaller plate, and treating it as the latter is what made the tracker twitch.

When a hand crosses a screen the detector does not fail — it succeeds, and finds the remaining visible fragment. Fitting a quad to that fragment collapses the tracked geometry onto it, so the composited insert shrinks and snaps back as the hand passes. On the project's laptop clip that left a wedge of raw plate exposed for the whole pass, and it is not something the matte can repair: the matte can only clip content that the warp actually placed there.

A matched detection is therefore checked against the track's own recent behaviour, and one whose plate area has collapsed is discarded rather than fitted. The track holds its geometry, stays open, and the skipped frames are bridged by the same linear interpolation that already covered detection dropouts.

Two expiries, because two different things are happening​

ExpiryGovernsDefault
MaxGapFramesA plate that has genuinely VANISHED — no detection at all3
MaxOccludedFramesA plate still detected every frame, just partially45

The distinction is available for free: an occluded plate is still producing a detection. Past MaxOccludedFrames the smaller reading is more likely to be what the plate now is than a very long occlusion, so it is accepted and the baseline re-establishes — which avoids both holding a stale geometry forever and splitting the track mid-shot.

Matching survives the hold​

A track held through an occlusion goes stale as a match target. Matching therefore runs against a constant-velocity prediction of the track's centroid, with a jump allowance that grows with the number of frames held. Without this a plate occluded while moving falls outside the centroid gate and is adopted by a brand-new track — which is exactly how the first version of this fix cut a handheld phone's single track in two.

Three things that had to be measured, and were wrong at first​

This is the part worth reading before changing any of it. Each of these was a plausible-looking choice that real footage falsified.

1. The test reads the STRICT CORE area, not the grown component. Hysteresis growth (see Tracked screen inserts (Phase 5)) intermittently bridges a plate into green scenery behind it. On the phone clip that makes the grown area flip between 0.19 and 0.35 frame to frame, while the core drifts smoothly by under 2%. Any threshold on the grown area is deaf or hysterical; the core is a steady signal.

2. The threshold is derived per track, not fixed. What counts as an abnormal dip is a property of the footage. Against a rolling baseline, the unoccluded handheld phone never drops below 0.980, while a hand crossing the laptop's screen takes it to 0.814 — but a locked-off shot of a bright plate holds to a fraction of a percent, and a fixed ratio suiting one is wrong for the others. Each track measures its own robust spread (median absolute relative step — median, so the very dips being hunted cannot inflate the yardstick used to find them) and sets its threshold at OcclusionSensitivity deviations below it, bounded by MinOcclusionDropMargin / MaxOcclusionDropMargin.

3. The baseline is a THREE-sample median, not ten. A longer window lags a plate genuinely moving away from camera until the lag itself looks like a collapse — a steadily shrinking plate had its track cut short until this was narrowed. Three is enough to shrug off a single bad frame, which is all the baseline needs; the spread is still estimated over a longer window, where more samples genuinely help.

Asymmetry is the whole trick​

Every part of this test is one-sided, and deliberately so: occlusion can only ever remove plate area. A plate moving toward the camera grows and is never suspected. One moving away shrinks gradually, which the short rolling baseline follows. Only an abrupt one-sided collapse looks like something passing in front.

Config reference​

VideoAnalyzeStepConfigDefaultMeaning
SolveInsertChromaKeytrueGrid-search each plate's colorkey parameters. Runs over already-sampled grid frames and is scored against the tracker's own segmentation, so it costs no extra decoding and needs no clip-specific configuration
VideoCompileStepConfigDefaultMeaning
EnableInsertMattetrueClip each insert to the plate's per-frame silhouette. A track with no solved key (or one that scored too low) composites with the plain quad mask regardless, so this can only improve a composite or leave it unchanged
ChromaQuadTracker.OptionsDefaultMeaning
OcclusionSensitivity6.0Robust deviations below a track's own behaviour before a core-area dip reads as occlusion
MinOcclusionDropMargin0.03Smallest dip that may ever count as occlusion — below this is sensor and compression noise
MaxOcclusionDropMargin0.30Largest dip a track may demand before believing an occlusion
MaxOccludedFrames45How long geometry is held before the smaller reading is accepted as the truth

Real-footage results​

CaseBeforeAfter
Phone's rounded bezel corners and camera notchpainted over by a filled quadfollowed exactly; notch cut out
Hand crossing the laptop screenpainted over; insert collapsed onto the visible fragment, exposing a wedge of platehand passes IN FRONT; geometry held through the pass
Residual plate inside the composite (laptop occlusion frame)2066 px545 px
Green rim tracing the handpresentremoved by the spill term plus despill; a trace remains at the blend limit

Limits, stated​

  • The blend limit is physical. A pixel that is genuinely half plate and half occluder cannot be made entirely correct; it can be absorbed into the matte or despilled, and both are done, but a faint trace can remain on a fast-moving edge.
  • The key is solved on ungraded pixels but applied to the base at the insert stage. A colour grade applied earlier in the chain shifts the plate's colour out from under it. EnableColorGrade is off by default, so this is latent rather than live; keying the original input instead is not possible because only the base is time-aligned with the output after the cut.
  • Hysteresis growth bridging a plate into green scenery was only worked around here at first; it is now fixed at the source by purity-gated, self-tuned growth — see Self-tuned growth and honest confidence.
  • Occlusion recovery is a hold, not an estimate. Geometry is interpolated across the occluded frames from the good frames either side; a plate that moves non-linearly while fully occluded will be slightly wrong in the middle of the pass.

Self-tuned growth and honest confidence​

What went wrong on the handheld phone clip​

A tracking test on greenscreen-phone.mp4 — a phone held almost still in front of a lawn — composited an insert that swung far outside the phone, reported Moving motion and 0.71 confidence, and painted over the phone's rounded corners. One cause produced all three:

  • The permissive grow mask is a fixed "green dominates by ~10%" test. The lawn reads about (70, 93, 26) and passes it. Wherever the coarse grid thinned the bezel to nothing, hysteresis growth flooded the plate ((0, 230, 0)) into the grass, and the extreme-point corners jumped onto the lawn on roughly half the frames — bottom-right corner at (0.68, 0.68) on one frame, (0.99, 0.97) on the next.
  • Every confidence term was computed one frame at a time, and each leaked frame was a tidy, well-filled quad, so nothing in the score could see the flicker.
  • The chroma-key solver scores candidate keys against the tracker's own segmentation. Half of that segmentation was lawn, so no key agreed, none was stored, and the compile step fell back to the plain quad mask — no silhouette matte, rounded corners painted over.
  • Separately, motion was the frame-to-frame centroid path length; one-grid-pixel corner jitter at 12 fps alone "travels" twice the Moving bar.

Purity-gated growth​

Growth now admits a pixel only if its chroma purity — how far the target channel rises above the strongest other channel, relative to the target, (g − max(r, b)) / g for green — reaches a fraction of the seed's own median purity. Purity is brightness-independent, so a dim region of a plate scores like its bright regions while dull foliage scores far lower. Measured on the real fixtures: the phone's lawn reads purity 0.1–0.3 against a plate of 1.0; the dim laptop plate's core reads 0.63 and its genuinely dim regions 0.5–0.6.

Searching the bar per video (ChromaTrackerTuner)​

What fraction of the seed's purity to require is a property of the footage, so it is not a constant: before tracking, ChromaTrackerTuner tries 1.5 (effectively no growth), 0.8, 0.65, 0.5, 0.4, 0.3, 0.2 and 0 on up to four blocks of 18 consecutive frames spread across the clip, and keeps the best by

score = detectionRate × edgeContrast × temporalConsistency

  • edgeContrast — along each fitted edge (10%–90% of its length, so a rounded corner's arc does not count against it), the purest pixel within 4 grid px inside minus the least pure within 4 px outside, normalized by the purity at the quad's centre. Only a fit whose edges sit on the real plate boundary has contrast across them: a fit that leaked into scenery has scenery on both sides of its far edge, one that stopped short inside a dim plate has plate on both sides.
  • temporalConsistency — the share of frames whose corners all sit within 4% of the quad's diagonal of their ±6-frame temporal median.

Ties go to the more conservative candidate. Measured on the five real fixtures:

ClipOld growth (bar 0)Under-growth (bar 1.5)Chosen barChosen scoreTrack confidence
greenscreen-phone.mp4 (lawn)0.42 (consistency 0.65)0.790.40.860.80 (was 0.71, wrong)
greenscreen-laptop.mp4 (dim)0.890.470.40.890.92 (was 0.64)
greenscreen-monitor-typing.mp41.001.001.51.000.97
c1-laptop-barrel30.mp40.850.440.50.870.91
c2-laptop-barrel48.mp40.770.470.50.850.87

The winning parameters and every candidate's score and measurements are stored on the track as VideoInsertRegionTrack.Tuning, carried into the timeline manifest, and shown in the editor, so a track can be explained from the artifact alone. ChromaQuadTracker.Options.GrowExcessFraction pins the bar and skips the search.

On the phone clip the result is a plate that stays on the screen (corners move by about 1% of the frame between frames instead of 30%), motion Slow, and a solved key (#00E200, 97% agreement), so the insert is matted to the phone's rounded corners and notch again.

Honest confidence​

Confidence is coverage × meanFillScore × meanEdgeContrast × consistency. The two new terms are what can see a wrong fit: with the bar pinned to the old behaviour the phone track now reads below the MinInsertConfidence floor of 0.5 and is refused rather than composited, while the correctly fitted dim laptop plates are no longer docked for needing growth.

Motion is classified from the smoothed track over half-second strides, where back-and-forth jitter cancels and genuine travel accumulates.

Tracking quality on the customer clips​

A customer's composited reels showed three things: the tracked screen jittered (worst on the finger-swipe phone clip and the turning tablet), raw green flashed through on some frames — a wedge down the phone while a finger crossed it at ~0:11, another down the tablet mid-turn at ~0:16, and the reel's last five seconds entirely green — and on a laptop/monitor reel the inserted UI was cropped at the plate's edges ("Dashboard" read "shboard"). Five separate causes, each measured and fixed separately; one of them was not a tracking problem at all.

The instrument​

InsertTrackingQualityHarness (tests project) runs the analyze step's own tracking pass over a directory of clips, composites a coordinate card into every track with the compile step's own filter builder and real ffmpeg, and reports per clip. It is skipped unless REELBOLT_INSERT_CLIPS names a directory; REELBOLT_INSERT_CLIP_LIST, REELBOLT_INSERT_LABEL, REELBOLT_INSERT_BASELINE=1 (every change below switched off where it can be), REELBOLT_INSERT_OPTS=Name=value,... (tracker option overrides, for sweeps), REELBOLT_INSERT_SAVE_FRAME=clip:frame,..., REELBOLT_INSERT_CONTENT/REELBOLT_INSERT_FIT (real content instead of the card) and REELBOLT_INSERT_DEBUG=1 (which detections the outlier pass replaces) steer it.

MetricWhat it measures
Residual greenPer output frame, the strict-plate-green pixels left in the composite as a fraction of those in the source frame; frames over 2% are a visible flash, over 50% a raw plate. A "visible green" variant also counts the dim, reflection-covered plate the strict test misses
Reference error / slipEvery tracked frame's edges re-fitted on the SOURCE frame at native resolution, seeded with the track's own corners, on frames where all four reference edges sit on a clean plate/bezel boundary (edge contrast ≥ 0.6). Error is the corner distance to that reference; slip is its frame-to-frame change — content sliding on the glass, which is what the eye reads as jitter, and which also exposes smoothing lag on fast moves
Content extentThe card encodes its own x in red and y in blue, so the visible content's range says how much of its frame survives on the plate
Content-frame offsetWhere the content's own corners land relative to the tracked plate under the old padded overscan (below)

A quadratic-residual "jitter" was tried first and rejected as a metric: generated footage moves in a staircase of held and doubled frames, and the content SHOULD follow that staircase — the whip clip read 4.25 px of "jitter" against 1.05 px of slip.

Cause 1: every insert ran one frame ahead of its plate​

perspective's in variable is 1-based — ffmpeg evaluates it as the link's frame count plus one — and the corner expressions were written with keyframe 0 at in = 0. So every insert used the NEXT frame's geometry. On a still plate that is invisible; on a moving one it is one frame's travel of misregistration, a strip of raw plate along the trailing edge and content sliding against the glass on every fast move. Verified directly (Perspective_in_variable_is_one_based...: a keyframe-0 edge at x=100 rendered at x=150, half way to keyframe 1). Keyframes are now written at in = k + 1 (ScreenInsertFilterBuilder.PerspectiveFirstFrame, both the flat and the mesh builder). On the whip clip this alone took residual green from p95 3.76% to 0.07% and its flashing frames from 10 to 5 (the rest are the whip's own first frames, see "What remains").

Cause 2: a finger over a corner, a reflection along an edge​

The per-frame edge-line fit is seeded from the plate's extreme-point corners, and a COVERED corner is not where the plate's corner is. On the swipe clip a fingertip over the lower-right corner dragged that corner up to 40 px along the finger; the right edge was fitted between the true top-right corner and the fingertip, and a green wedge opened for two thirds of a second. The tablet did the same along the boundary of a green-tinted window reflection on its glass. The occlusion test in track assembly cannot see either: the core area falls only a few percent while a finger enters, and a reflection costs no area at all.

ChromaEdgeTemporalRefiner re-locates every frame's edges from its NEIGHBOUR's. A plate moves little between frames, so the previous frame's edge is a far better seed than this frame's own corners. Per edge, candidates are scored by ChromaEdgeLineFitter.EdgeContrastScore — dominance just inside the line minus just outside it, over the plate's level, along its middle 80%: a real screen edge has plate inside and a near-black bezel outside and scores near 1; a line on a finger has plate on both sides along most of its length, and a line along the tablet's reflection has dim plate (dominance ~25 against a level of ~140) outside it.

CandidateWhy it is offered
The detected edgeWins every tie (PreferenceMargin 0.05), so a clean frame is left exactly as it was
A re-fit seeded on the neighbour's edgeProbes that start on the finger or the reflection are skipped as holes; the robust fit ignores the minority that cross the wrong boundary
On a frame bridged across an occlusion: the detection the occlusion test set aside, and a re-fit seeded on itIts uncovered edges are still the plate's real edges (FrameQuad.SetAside)

Two rules keep it from undoing what the occlusion hold protects. A candidate may not move an edge net inward by more than 5% of the plate's extent across it (InwardToleranceFraction): something in front of a plate can only remove plate area, so a line further in is an occluder's boundary — "net", because the corruption repaired here TILTS an edge, inward at one end and outward at the other. And corners move at most 35% of the diagonal (the reflection needed 21%). The pass runs forward then backward, so a track whose first frames are the corrupted ones is repaired from the clean frames after them; it covers flat tracks only, and treats a frame of a track that reads flat overall (median curvature under the compile step's 0.01) as flat even when its own profile read a bow (motion blur).

Cause 3: plates with nothing composited into them​

Three ways a plate ends up on screen with no insert, all of them seen on the customer's reels:

  • The track was dropped. The reel's last five seconds were not an untracked clip: its fourth phone's track existed (confidence 0.62) and was dropped max_inserts_exceeded — MaxInserts defaults to 3 and the reel had four clips.
  • Frames the tracker deliberately leaves out — a phone turning through glare (TrimWeakEnds), the first frames of a whip into place, a plate seen too briefly to become a track. On the man-with-phone clip that is frames 19-21 and 25, at full raw green.
  • Output time outside the insert's window, e.g. a track cut into two kept portions.

The analyze step now records, per track, UncoveredPlateSpans: every run of frames where the plate is visible (strict plate pixels covering a quarter of MinInsertRegionAreaRatio) but no track covers it, attached to the nearest track in time, with the bounding box of the plate pixels seen over the run. The compile step's InsertPlateGuardPlanner turns every such window — and every window of a track that is not being filled — into a plate guard: an ordinary insert whose content is a dark neutral screen (0x101010, an in-graph lavfi colour), matted with the track's own solved key exactly as a real insert is, so it can only darken pixels that ARE the plate, with the plate's reflections carried over so it reads as a screen that is off. Geometry is the track's own quads inside its span, and the recorded box (grown 10% a side) over an uncovered span. A guard never overlaps the insert it accompanies. Each window is reported in the EDL (inserts.unfilledPlates: region, reason, output window, neutralized) and the compile adds one customer-facing C1 warning, insert_untracked_plate_visible. A track with no solved key, or far under the confidence floor (below half of MinInsertConfidence — it may not be a plate at all), is reported but not darkened.

To fill the reel's fourth phone rather than darken it, raise MaxInserts to 4.

Cause 4: the cropped laptop title was the configured fit​

The "Dashboard" → "shboard" reel was compiled with insertFit: "Cover", not Auto: Cover crops the content to the plate's shape, and the laptop plate's aspect is 1.40 against the UI's 1.78, so 10.5% came off each side — the sidebar and the start of the title. Re-rendered with the same tracks and content: Cover reproduces the crop exactly; Auto contains it (whole frame, dark bars top and bottom); Stretch fills the plate with the content squeezed 21% horizontally. Evidence crops are kept in docs/screenshots/tracking-2026-10/ (laptop-as-delivered-title.png, laptop-new-cover.png, laptop-new-auto.png, laptop-new-stretch.png). Use Auto or Contain for content that must be seen whole.

The overscan was a suspect too, and measured a real but smaller error. A matted insert's warp target is the plate grown by 10% about its vertex centroid, and the content used to be padded by the same factor (black) so its own frame landed back on the plate — exactly, only for a parallelogram. On a foreshortened plate the content's corners landed off the plate's, by a mean 2.8 px (max 4.9) on the laptop, 2.4 px (max 7.3) on the tablet, and up to 86 px on the whip's degenerate first quads — a strip of black pad inside the plate, or content cropped. Smeared overscan (ResolvedScreenInsert.ContentKeyframes) replaces it for flat matted inserts: the content's frame is warped onto the tracked plate itself, without the 1px border, and perspective's edge clamping smears its outermost pixels across the band, where the matte keeps them to the silhouette. The content's corners are on the plate's by construction, and the band between the tracked edge and the true edge shows the content's own edge colour instead of black. A mesh insert keeps the pad.

Cause 5: smoothing — measured, left alone​

With the corrupted frames repaired, the remaining slip on the moving clips (0.5-1.05 px) is near the reference's own noise. A sweep of the smoothing found no setting better on every clip, so the defaults stand:

Settingswipetabletwhipman-phonephone
Shipped: window 3, motion gate 2× noise0.900.541.051.050.37
Gate 1×1.180.581.421.100.32
Gate 4×0.780.550.801.240.41
Window 50.790.641.110.940.50
No smoothing1.430.681.761.270.29

(slip RMS, source px.) Disabling the outlier rejection was also tried: it lowered slip on the swipe (0.85) and whip (0.91) clips — the rejection does smooth away some real frame-hold cadence — but let three swipe frames flash green (up to 6.7%), so it stays.

Before and after​

Before is the shipped tracker and compile; after is all of the above. Residual green: mean / p95 / frames over 2% / frames over 50%. Reference: mean / p95 / max corner error, and slip RMS, in source pixels.

ClipResidual beforeResidual afterReference beforeReference after
Finger swipe (1344x768)0.61% / 5.31% / 16 / 00.02% / 0.07% / 0 / 02.28 / 9.35 / 48.0, slip 1.670.78 / 1.94 / 3.8, slip 0.90
Tablet turning by a window0.39% / 2.81% / 12 / 00.02% / 0.18% / 0 / 00.77 / 1.81 / 2.6, slip 0.54 ¹0.77 / 1.81 / 2.6, slip 0.54 ¹
Phone whipped up (768x1344)2.10% / 3.76% / 10 / 11.82% / 0.01% / 5 / 1 ²0.78 / 1.89 / 4.3, slip 1.050.78 / 1.89 / 4.3, slip 1.05
Man turning a phone toward camera2.95% / 0.07% / 4 / 40.44% / 0.00% / 1 / 1 ³0.98 / 1.93 / 2.7, slip 1.050.98 / 1.93 / 2.7, slip 1.05
greenscreen-phone.mp4 (control)0.26% / 0.57% / 0 / 00.26% / 0.57% / 0 / 00.70 / 1.24 / 2.1, slip 0.37unchanged
Laptop on a desk (control)0.00% / 0.00% / 0 / 00.00% / 0.00% / 0 / 00.60 / 1.83 / 2.5, slip 0.38unchanged
Monitor (control)0.00% / 0.00% / 0 / 00.00% / 0.00% / 0 / 00.67 / 0.99 / 1.0, slip 0.01unchanged

¹ The reference itself fails on the tablet's reflection frames, so it covers only the 50 clean ones, where nothing changed; the residual column is where the reflection fix shows. ² The five frames left are the whip's first, see below. ³ One frame, see below. Before/after composites of the swipe (frame 116), the tablet (frame 101) and the man-with-phone turn (frame 20) are in docs/screenshots/tracking-2026-10/.

What remains​

  • The first ~0.3 s of a whip into place. Frames 40-46 of the whip clip — the phone at 45°, heavily motion-blurred, entering the frame — still fit a wrong quad, and 23-44% of the plate shows on four of them. The refiner cannot help (its seed is the equally wrong neighbour), and the guard does not overlap a running insert.
  • A plate washed out by glare keys only partially, so a guard darkens it only partially (man-with-phone frame 25, 59% left).
  • A plate in a clip with no track at all (every appearance shorter than MinInsertRegionSeconds) has no key and no geometry to guard with, and still shows green.
  • Cover crops by design.

Memory: every branch starts at zero​

A four-clip landscape compile with tracked screen inserts was SIGKILLed (ffmpeg encode failed (exitCode=137)) on a host with 44 GB free and no container limit: ffmpeg itself grew to tens of gigabytes. The cause was the shape of a matted insert's filtergraph, not the number of clips.

Why it buffered. A matted insert's alpha is keyed off the base picture (see Silhouette matting), so the base is one split with a branch per matte, despill and reflection layer. The insert's content and quad-mask branches used to be moved onto the output clock with a timestamp shift, setpts=PTS-STARTPTS+start/TB, so their FIRST frame sat at the insert's start. Every framesync filter (overlay, alphamerge, blend) must see a frame on each input before it can emit anything, so the overlay could not pass output frame 0 until the insert branch produced its first frame — and that frame needed the matte at start, which needed the base decoded up to start. split hands every frame to every output whether or not it is consumed, so the whole programme before the insert's start was queued, full resolution, in RGB, once per split branch. Memory grew with how LATE an insert started, times the number of matted branches; a four-clip composite's last insert starts 20 s in.

The fix (ScreenInsertFilterBuilder.DelayToOutputStart). A matted insert's branches are no longer shifted; they are PADDED from t=0 with black frames (tpad=start=N:color=black, N = the shift in whole frames), which tpad generates without pulling its input. Every branch now carries a frame for every output frame and advances in lockstep with the base. The pad is black, which on the mask is zero coverage, and the overlay's enable window gates it anyway; the pad follows perspective, so its per-frame in still counts from the insert's first frame; the mesh path pads each cell after its own warp. Unmatted inserts keep the shift — nothing they wait on comes from the base. A split of one content file for several plates was deliberately NOT introduced: split branches consumed at different times are exactly what buffers, and a separate -i per insert costs only a decoder.

Compile (1920x1080, preset slow)Peak RSS beforePeak RSS after
Four clips, 4 inserts + 3 plate guards> 20 GB (killed by the measuring harness; SIGKILLed in production)7.6 GB
Two clips (man with phone, finger swipe), 2 inserts + 2 guards14.7 GB4.7 GB
Same two-clip graph without the encoder (framemd5)13.6 GB4.1 GB

The two-clip graph's decoded output is bit-identical before and after (all 733 video and audio framemd5 lines). ScreenInsertMemoryShapeTests pins the shape (no matted branch may contain a PTS-STARTPTS+ shift; every one is padded after its warp) and measures a small real-ffmpeg case: a matted insert 20 s into a 640x360 programme peaks at ~190 MB padded, ~900 MB shifted.

What memory still scales with. Not with how late an insert starts any more, but still with the number of matted inserts and plate guards: each is a few dozen full-resolution filter links (matte, despill band, reflection layer, warp, composite), and ffmpeg keeps a small pool of frames per link — about 0.6-0.9 GB per matted insert at 1080p (measured: 0.4 GB for the bare four-clip concat, 3.9 GB with four matted chains, 6.75 GB with seven). Restricting each chain to its plate's bounding box would shrink that roughly by the box's share of the frame; it is not done.

Measuring a compile. InsertCompileMemoryHarness (tests project, skipped unless REELBOLT_INSERT_MEMORY_DIR names a directory holding the analysis artifact as analysis.json, each source as <projectFileId>.mp4 and the insert content as content.mp4) runs the real VideoCompileStepExecutor against them and, instead of encoding, copies the scratch space and the encode's argv into capture-<label>/ so the exact ffmpeg invocation can be run under /usr/bin/time -v. REELBOLT_INSERT_MEMORY_SOURCES=0,1 compiles a subset of the clips.

A clip's last frame and the next clip's first​

On a multi-clip compile with inserts, the FIRST frame of every clip after the first showed its raw green plate — neither the insert nor the dark plate guard covered it (16% and 21% of the frame on the customer's two-clip composites); single-clip compiles were clean.

Generated clips carry fewer video frames than their container says: 158 frames at 24 fps (6.583 s) in a 6.592 s file, the audio being longer. A span kept to a clip's end is snapped to the container's duration — 159 frames, 6.625 s — so the output timeline put the next clip at 6.625 s. The segment delivered only 158 frames, concat started the next clip at the audio's 6.592 s, and that first frame fell before every insert's and guard's enable='between(t,6.625,...)' window. Worse, every later frame of that clip paired with the previous frame's geometry. The timing in the filters was right; the picture disagreed with the timeline.

Each segment of the segmented (concat) encode is now held to EXACTLY the length the timeline gives it (VideoCompileStepExecutor.SegmentLengthSuffix): tpad=stop_mode=clone:stop_duration=<len>,trim=end=<len> — any shortfall is filled with the clip's own last frame, and a segment that really has every frame is untouched. On the customer's two-clip composite, frame 158 went from 16.5% raw green to none, and the whole four-clip composite has no frame over 0.15%. SegmentLengthBoundaryTests reproduces it with real ffmpeg on a clip shaped like the generated ones (and shows the unpadded graph starting the next clip a frame early).

What remains at the 0.15% level is a different thing: on the finger-swipe clip around 7.7 s, a hand crossing the phone's lower-right edge pulls the tracked edge inward for a few frames, leaving a thin strip of plate beside the hand — a tracking limit of the kind described in Tracking quality on the customer clips, not timing.

Inserting an existing video​

An insert's content used to be only something the motion-graphics agent rendered in the same run — the compile step forces such keys under the execution's own outputFiles/ prefix. There was no way to say "put my clip on the phone", so a workflow asking for that had to list the clip as a second VideoAnalyze source. That offered its shots to the story editor as footage, the editor kept them, and the finished video was the phone take with the entire promo appended after it, while the planner rendered a look-alike promo of its own for the screen.

Two paths now exist:

  • Deterministic: VideoCompileStepConfig.InsertContentProjectFileId names a video/* or image/* project file. With EnableInserts, it is composited into every offered region that clears MinInsertConfidence, most confident first, up to MaxInserts, skipping regions a plan already filled. No agent and no GraphicsPlan are needed — a tracking test is VideoAnalyze (the plate footage only, DetectInsertRegions: true) → VideoCompile. Still images are looped for the insert's whole window.
  • Planned: a ScreenInsert's renderedAssetStorageKey may be a project file's id (or its exact storage key). It is matched against this project's own file list; anything else is treated as an agent render and re-anchored exactly as before. The planner's prompt now tells it to use the user's file instead of authoring a substitute, and the story editor's prompt tells it that a clip the request describes as content to place inside the picture is not footage to keep.

InsertFit (Auto | Stretch | Contain | Cover) shapes the content to the plate's measured physical aspect before the warp maps its whole frame onto the plate. Auto always stretches an agent render (it was authored for the plate, so it is unchanged) and a project file whose aspect is within 15% of the plate's, and otherwise contains — a landscape promo plays letterboxed on a portrait phone instead of squashed. The EDL's appliedInserts records each insert's contentOrigin (plan/config), contentProjectFileId and contentFit.

Different content per screen (InsertContents)​

InsertContentProjectFileId puts ONE file on every plate, so a portrait UI for the phones and a landscape UI for the laptop in the same edit took two compile passes. InsertContents (a list of VideoInsertContentAssignment(ProjectFileId, SourceIndex?, RegionId?, Fit?)) assigns content per plate on the deterministic path. For each offered screen plate (feature-tracked objects excluded, as before), PickInsertContent takes the most specific matching entry — one naming the plate's RegionId outranks one naming its SourceIndex (the analyze step's 0-based clip), which outranks one naming neither; list order breaks ties — and falls back to InsertContentProjectFileId; a plate neither covers gets no deterministic content. An entry's Fit overrides InsertFit for that content. Every other rule (offered ids, confidence, MaxInserts, plan-filled plates first) is the shared loop's, unchanged. Null/empty is byte-identical to the single-content behaviour.

Dips to black in insert content (InsertContentDips)​

Screen recordings dip to black between screens (one measured at 2.8–3.6 s, 5.7–6.0 s, 8.6–9.1 s and 11.3 s), and on a tracked phone that reads as the phone switching off. For every project-file video used as insert content (renders made for the plate and still images are never touched), the compile runs ffmpeg's blackdetect once per file (d=0.08:pic_th=0.98:pix_th=0.10 — a frame with 98% of its pixels under 10% luma, for at least 80 ms; parsed by ParseBlackDetect) and, per InsertContentDips (InsertContentDipMode, append-only):

ValueWhat is composited
Freeze (default)A copy where each dip's frames are dropped and fps=…:start_time=0 refills the gap by repeating the last good frame — the content keeps its length and timing; a dip at the very start shows the first good frame, one running to the end is held by tpad
CutA copy with every dip removed (select + setpts) — the content gets shorter
KeepThe file as it is; nothing is detected (no extra ffmpeg call)

The copy is re-encoded once (libx264 -crf 16) and reused for every plate showing that file (BuildInsertDipFilter). Content that is at least 90% black is left alone (content_mostly_black); a failed detection or re-encode composites the original (bridge_failed) — never a failed compile. Each applied insert whose content had dips carries, in the EDL and output summary:

"contentDips": { "treatment": "Freeze", "intervals": [ { "startSec": 2.8, "endSec": 3.6 } ],
"totalSec": 0.8, "applied": true }

(absent when no dip was found, so an insert with clean content keeps its exact shape). Because the default is on, a compile with insert content now makes one blackdetect pass per content file.

Tracking in the editor​

The timeline manifest's insert items carry, besides the source-space quadKeyframes the EditTimeline v2 seeder reads, outputQuadKeyframes — the corners the encode actually warped to, in output timeline seconds and 0..1 of the output picture, after the cut, overscan and any letterbox fit — plus matted and tuning. The program monitor draws those inside the video's displayed rectangle. It previously multiplied source-space corners by the whole monitor canvas, so a portrait video pillarboxed in the landscape monitor had its tracked corners drawn out in the black bars, and it labelled any plate with non-zero measured curvature "mesh" even when it was composited with a planar pin.

Tracking objects​

VideoAnalyzeStepConfig.TrackObjects finds and follows arbitrary flat objects — a licence plate, a sign, a screen with no green plate — with a markerless planar tracker (Services/Video/Tracking/, pure C#, no new dependency). The chroma tracker above needs a uniform-colour plate; this one needs only texture.

Pipeline, per source (ObjectTrackingService)​

  1. Frames. One grayscale pass (IFrameGridSampler.SampleGrayAsync, bicubic — the tracker needs gradients, which an area downscale blurs) at ObjectTrackSampleFps (default 15), longest side ObjectTrackMaxSide (default 960).
  2. Seeds. A target with a Region (normalized corners at AtSec) seeds exactly there with no model call. A target with only a Label is FOUND: on frames every ObjectDetectEverySec (default 1 s; past MaxObjectDetectFrames the lookups are spread evenly over the whole clip) the Vision provider (IObjectLocator/VisionObjectLocator) returns boxes on a 0..1000 integer grid. Each box is then refined on a zoomed crop around it (CropAround: 2.5x the box, at least 12% of the frame, re-extracted at 768 px with ExtractCroppedKeyframeAsync), because vision models place boxes loosely on a full frame and much more tightly when the object fills the image. The model contributes only seed boxes; the label is the author's text, sanitized to one short line before it enters the prompt.
  3. De-duplication. Seeds are processed placed-first then in time order; a seed whose box overlaps (IoU > 0.3) the same label's existing track at that frame is the object that track already follows and is skipped. This is what turns "looked every second" into one track per object rather than one per lookup, while still catching objects that enter later or that a track lost.
  4. Tuning, once per label (PlanarTrackerTuner): motion model (homography / affine / similarity) × LK window radius (5, 7, 10, 13). Each candidate tracks the seed up to 16 frames out, then tracks the quad it ended on back to the seed; score is coverage × meanConfidence × exp(−roundTripError / 0.03) with the error as a fraction of the seed quad's diagonal. Pyramid depth was searched too at first and dropped: depths 3 and 4 gave bit-identical tracks for every candidate.
  5. Tracking (PlanarTracker), forward and backward from the seed, per frame:
    • Shi-Tomasi features in the quad grown by DetectionMargin (15%), so the object's own outline corners count; background features caught in the margin move differently and RANSAC drops them.
    • Pyramidal Lucas-Kanade (Bouguet) frame to frame, keeping only points whose backward flow returns within 1 px (forward-backward check).
    • RANSAC fit of the tuned model (deterministic seed, Hartley-normalized least squares).
    • Drift control: the chained estimate is used only as the starting guess for registering the SEED frame's features directly onto the current frame; when that registration agrees, it replaces the chain, so error does not accumulate frame over frame. It falls back to the chain when the target has changed too much to register.
    • Sanity (convex quad, per-frame area change within 35%, centre near the frame) and an appearance gate — NCC of the rectified quad against the seed and a slowly updated recent template. Below 0.3 for more than 3 frames ends the track instead of letting it slide onto the background.
    • Features are replenished inside the quad when fewer than half survive, mapped back to seed space so they join the reference registration.
  6. Output. Each run becomes a VideoInsertRegionTrack with Method = "feature", its Label, 3-frame-smoothed keyframes, Confidence = coverage × mean per-frame confidence (per-frame: RANSAC inlier ratio × appearance), and the Tuning record. Tracks join the source's insert regions after the chroma plates, continuing their r{n} ids, so everything that consumes an insert region — screen inserts included — works on them. The view marks them "kind": "object" with their label.

meta.objectTracking (applied, degraded, reason, tracks, visionCalls) appears only when TrackObjects is configured. No Vision provider, a failed lookup, or a tracker exception all degrade with a reason; none fails the step.

Measured​

On a synthetic clip with exact ground truth — a licence-plate graphic composited onto the real broll-04-servers.mp4 with ffmpeg's perspective filter, growing 35%, rotating and skewing over a camera move in the other direction, 135 frames at 960x540 — the tuned tracker followed every frame with a median corner error of 4.4 px at 1080p (p95 8.5, max 18). Untuned (homography, 21 px window) it drifted to 49 px by the end, which is what the round-trip search exists to catch.

Censoring tracked objects​

VideoCompileStepConfig.CensorLabels obscures every feature-tracked region whose label matches (case-insensitive; "*" = all), CensorStyle Blur | Pixelate | Fill.

  • Every portion. OutputTimeline.MapWindowToOutputAll maps a track into EVERY output stretch its source window survives into. Inserts use the first stretch only; for privacy that would leave an object visible after a cut.
  • No confidence gate. A doubtful track is still obscured — over-obscuring is the safe failure.
  • Padding and hold. Each region gets a margin of CensorPadding (default 0.25) times its LONGER side on every side — a licence plate is ~4x wider than tall, and a margin proportional to its own height added only a few pixels where a vision box a little off vertically needed the most cover (on the car-meetup clip a box 16 px too high left the bottom row of characters readable). Each track's first/last position is held CensorHoldSec (default 0.4 s) outward.
  • Rendering (CensorFilterBuilder): the picture is split and the copy obscured ONCE (Gaussian blur sigma 1.5% of the canvas width; 2.5%-of-width mosaic; or solid black). One combined mask is drawn on a quarter-resolution black canvas: per region and per one-second chunk of its track, a white box as large as the region's largest padded extent in that chunk, moved every frame by an overlay x/y expression in output time (a balanced if(lt(t,…)) tree, since ffmpeg rejects deeply nested expressions). The mask is scaled up, alpha-merged with the obscured copy and overlaid back. Twelve plates on 16 s of 1080p: 64 s when every region obscured the whole frame, 9 s now.
  • Why boxes, not the insert warp. The first version pinned a white plate onto each region with perspective, exactly like a screen insert. ffmpeg's perspective precomputes fixed-point coordinate maps that overflow when a full-frame plate is squeezed ~20x onto a plate-sized quad; the "mask" covered everything above-left of each plate. Screen inserts never hit this because screens are large. Axis-aligned boxes over the padded quad need no warp, and erring large is the right bias for privacy.
  • Stage order. It is the last picture stage before the program fade, so anything composited onto the object is obscured too; on the segmented path it runs on the program body, before generated cold opens/end cards are concatenated, so its body-timeline windows stay aligned.
  • Fail closed. A compile with CensorLabels fails with CENSOR_UNVERIFIED instead of producing a video that may show what it was meant to hide when: the analyze step tracked no objects at all; any censor label (other than *) is not one the analyze step looked for in every source (Provenance.ObjectTrackingLabels, compared case-insensitively — "licence plate" vs "license plate" is a mismatch, deliberately); or object tracking DEGRADED — no Vision provider, a failed or unreadable lookup (a refusal or truncated reply is a failure, never "nothing here"), more than 16 instances in one lookup, a tracker error, or sightings left untracked because MaxObjectTracks was reached. A clean run that genuinely found nothing compiles, and the EDL's censor.unmatchedLabels (also reported as a progress message) names each label that matched nothing. CensorLabels with Mode = StreamCopy is a hard failure (CENSOR_REQUIRES_REENCODE). Outputs whose objectTracking, vision or insertTracking stage degraded are never written to the step cache (StepResultCache.ReportsDegradation), so a fixed provider is not masked by a cached failure. Transcription is deliberately excluded: it reports "degraded" for a source with no audio or an install with no speech-to-text provider, which are permanent facts of the input.
  • The timeline manifest carries a censor track whose items have outputQuadKeyframes (the tracked quad, unpadded); the editor draws them as red dashed outlines with label and confidence.

Measured on real footage​

vintage-cars-meetup.mp4 (Wikimedia Commons, "Renkontiĝo de ŝatantoj de malnovaj aŭtoj en Tjumeno (2022)", CC BY-SA 4.0): 16 s handheld, two parked cars, people walking in front of both plates. 36 vision lookups (DeepSeek Vision, via the plain-JSON fallback), 17 tracks; both plates are obscured in every sampled frame from 0 to 16 s. Over-detections (a turn signal, a patch of pavement) are obscured too; that is the intended bias.

KeepWholeSources​

VideoCompileStepConfig.KeepWholeSources ignores Decision and keeps every analyzed source whole, in source order (WholeSourcesDecision, built from the artifact's own shot ids and checked against every shot rather than the offered set). It is the no-agent path for compiles that add to footage rather than cut it: a tracking test, a censor pass, a screen insert of a project file.


The edit room​

StepType.EditRoom replaces the single AgentType.VideoStoryEditor decision step with a multi-agent deliberation: several editor-role seats plus a director converse in a live Microsoft.Agents.AI.Workflows group chat over the same bounded VideoAnalyze view a solo editor would see, and the director synthesizes their discussion into ONE schema-validated VideoEditDecisionOutput — the exact same schema and the exact same rushcut invariant (never a timestamp, only offered ids) AgentType.VideoStoryEditor already produces. VideoCompileStepExecutor needs zero changes to consume it: Decision just points at the EditRoom step's StepOrder instead of a solo VideoStoryEditor step's.

┌─────────────────┐ ┌──────────────────────────┐ ┌──────────────────┐
│ StepType. │ │ StepType.EditRoom │ │ StepType. │
│ VideoAnalyze │────▶│ several editor seats + │────▶│ VideoCompile │
│ │ │ a director, live group │ │ │
│ (unchanged) │ │ chat, then one synthesis │ │ (unchanged) │
│ │ │ call outside the chat │ │ │
└─────────────────┘ └──────────────────────────┘ └──────────────────┘
emits VideoEditDecisionOutput
(+ an additive "room" metadata
block) — identical shape to a
solo VideoStoryEditor step

Why a group chat, and why it's deterministic-scheduled​

Every seat sees the identical bounded view and can speak to any part of it — there is no information asymmetry between seats that would make an LLM-driven "who should speak next" routing decision meaningful. EditRoomGroupChatManager (WorkflowEngine/Agents/EditRoom/EditRoomGroupChatManager.cs — since the graphics room landed, a thin binding of the room-generic RoomGroupChatManager base, contributing only the edit room's [sgt]{n} offered-id regex; see The shared room infrastructure) is therefore a plain, deterministic round-robin scheduler over the configured seats, followed by the director: it makes zero model/network calls itself, only orchestrating which already-constructed AIAgent speaks next. Routing this through an LLM would double the room's cost for a decision that doesn't need intelligence.

The two agents​

  • The editor seats (EditRoomStepConfig.Seats, three by default — PacingEditor/ StoryEditor/CraftEditor, argued personas for rhythm/narrative/material-quality respectively) all resolve to the SAME built-in AgentType.VideoStoryEditor agent (its own seeded prompt/tools/provider) — no new AgentType enum member exists per seat, since only the persona differs, and personas are injected per-turn, not baked into separate agent definitions.
  • AgentType.VideoEditDirector is used TWICE, through two different code paths: once per-turn as a ROOM PARTICIPANT (free-form prose, moderates disagreement, ends the room by emitting the literal sentinel ROOM_DECIDED once satisfied), and once more, entirely OUTSIDE the group chat, for a single ordinary structured-output SYNTHESIS call (ReelBoltAgentBase.RunAsync, the same mechanism every other structured-output agent uses) that converts the room's transcript into the final VideoEditDecisionOutput. Same minimal read-only tool scope as VideoStoryEditor (no sandbox, no write/render tools — it only decides).

How a turn is built: RoomSeatAgent​

RoomSeatAgent (WorkflowEngine/Agents/Rooms/RoomSeatAgent.cs — named EditRoomSeatAgent until the graphics room landed; renamed unchanged since it was already fully room-agnostic) wraps each already-constructed inner AIAgent (built by EditRoomStepExecutor via the same chat-client- resolution path ReelBoltAgentBase.CreateAgentAsync uses) as a DelegatingAIAgent, overriding BOTH RunCoreAsync and RunCoreStreamingAsync — the group chat host always invokes participants through the STREAMING path, so a wrapper that only overrides the non-streaming one is silently bypassed (measured live against the real rc2 package). Three things happen on every turn:

  1. Per-turn sampling options are injected, since the group chat host always passes options == null to a participant — this wrapper is the only way to control temperature/reasoning-effort/MaxOutputTokens per turn (EditRoomStepConfig.Temperature/ DirectorTemperature/ReasoningEffort/MaxTurnTokens), via the same RawRepresentationFactory
    • OPENAI001 mechanism ReelBoltAgentBase.BuildChatOptions already uses for reasoning_effort.
  2. The seat's persona is appended as the LAST message, after the identical [system instructions][bounded-view opening message] prefix every seat/the director share — measured live as a ~2-3x latency win, since it lets that identical prefix keep hitting the backend's prompt-prefix cache. Putting per-seat identity earlier defeats the cache; this ordering must not be changed.
  3. A seat's turn throwing never aborts the room. Wrapped in try/catch: on failure, a synthetic "[{SeatName} had no input this round]" turn is recorded instead, and the room continues.

Scheduling and termination​

EditRoomGroupChatManager.SelectNextAgentAsync round-robins the editor seats for EditRoomStepConfig.Rounds full passes, then always returns the director:

editorCount = seats.Count
turn i: i < editorCount * Rounds → editors[i % editorCount]
otherwise → director

EditRoomStepConfig.Termination gates how the room can end EARLY, on top of the hard MaxTurns ceiling (mapped to GroupChatManager.MaximumIterationCount, clamped 2..20 — deliberately far below the framework's own default of 40, which is a multi-hour runaway on this backend, not a safety net):

ModeBehavior
SentinelOnlyEnds when the director's most recent turn contains the literal ROOM_DECIDED token.
ConvergedEnds when the offered-id-vocabulary mentions across the last MinConvergenceRounds consecutive editor rounds are identical — the seats have stopped proposing anything new.
SentinelOrConverged (default)Either check ends the room early.
FixedTurnsNeither check runs — only the MaxTurns ceiling ends the room.

Both checks are evaluated in ShouldTerminateAsync, called BEFORE any agent has spoken too (iteration 0, history = just the opening message) — the checks must not assume at least one turn has happened, and don't. UpdateHistoryAsync is a pure pass-through: it observes every new message for the executor's own transcript/progress bookkeeping but always returns the input history unchanged — returning anything else would corrupt the framework's canonical transcript, since this return value is NOT a per-turn view.

The executor: EditRoomStepExecutor​

Same never-throws, always-valid-JSON discipline as VideoAnalyzeStepExecutor/ VideoCompileStepExecutor (see The three-stage shape):

  1. Deserializes EditRoomConfigJson; resolves the bounded view via EditRoomStepConfig.View (an ExtractInputRef, reused verbatim from VideoAnalyzeStepConfig — the same Previous/Step resolution VideoCompileStepExecutor.ResolveDecisionJson already established for Decision/ GraphicsPlan/MusicPlan) and extracts the offered shot/silence/segment id vocabulary from it.
  2. Builds every seat's and the director's AIAgent, wraps each in RoomSeatAgent, and runs them through AgentWorkflowBuilder.CreateGroupChatBuilderWith(...).AddParticipants(...).Build() via InProcessExecution.RunStreamingAsync, bounded by EditRoomStepConfig.RoomTimeoutSeconds. The bounded view is sent as ONE opening chat message (not folded into the agent instructions), since that message becomes the shared prefix every seat's every turn hits the prompt-prefix cache against.
  3. Reads the transcript back off the WorkflowOutputEvent the workflow emits once it completes.
  4. Makes ONE standalone structured-output synthesis call (AgentType.VideoEditDirector.RunAsync, outside the group chat) with the view + rendered transcript, retried up to MaxSynthesisAttempts on an empty/unparseable result.
  5. Validates deterministically, never trusting the model: any Keep span whose FromId/ToId isn't in the offered-id set is dropped, and — when the view spans more than one source clip (see Multiple source clips) — so is any span whose FromId/ToId carry two different "src" indices, the exact shape VideoCompileStepExecutor rejects hard as MIXED_SOURCE_SPAN (dropping it here instead keeps the step's degrade-not-fail discipline, and the remaining single-clip spans still compile). Both are recorded in the output's room.droppedSpanCount; the mixed-source subset also in room.droppedMixedSourceSpanCount. A synthesized decision containing a mixed-source span first gets a retry with precise feedback (within MaxSynthesisAttempts) before the drop is accepted. The room charter prompt every seat (and the director's room-participant turns) runs under states the same one-hard-rule the solo VideoStoryEditor prompt's "Multiple source clips" section does — a single kept run's first and last id must come from the SAME clip — guarded by Room_charter_prompt_carries_the_multi_source_single_clip_span_rule.
  6. If the room failed outright, produced zero usable turns, the synthesis call never produced a usable decision, or every Keep span got dropped as unoffered — and FallbackToSoloEditor (default true) — falls back to ONE ordinary solo AgentType.VideoStoryEditor call: today's existing single-editor pipeline, unchanged. The step still completes successfully with a valid decision; room.degraded/room.degradeReason record what happened.
  7. Emits output_json = the validated VideoEditDecisionOutput, camelCase-serialized, plus an ADDITIVE sibling "room" object (seats, rounds, turn count, how the room terminated, whether it converged, synthesis attempts, dropped-span count, whether it degraded and why, the transcript artifact's storage key). This is safe because VideoCompileStepExecutor's own deserialization of VideoEditDecisionOutput uses no UnmappedMemberHandling.Disallow — the extra "room" key is silently ignored by that step exactly like any other consumer expecting the plain schema.

The transcript artifact — NEVER authoritative​

When EditRoomStepConfig.PersistTranscript (default true), the full, unabridged room transcript is uploaded as this step's ArtifactStorageKey, under the same projects/{projectId}/agentFiles/ video-analysis/{executionId}/step-{order}-room-transcript.json prefix convention VideoAnalyze/ VideoCompile already use for non-playable JSON artifacts (see Where artifacts live) — reachable through the same GET /api/v1/projects/{projectId}/ step-results/{stepResultId}/artifact endpoint. This transcript is free-form model prose and must never be parsed back into a decision. A seat or the director can say something that sounds like a timestamp in passing prose ("that pause feels like it's about three seconds") — harmless as commentary, but nothing downstream may ever try to extract a number from it. The only authoritative output of an EditRoom step is the synthesized, deterministically-validated VideoEditDecisionOutput on output_json.

Live progress and the transcript event​

Every completed turn reports a context.ReportProgressAsync stage/percent update (e.g. "PacingEditor is proposing a cut (turn 3/8)" / "Director is reviewing") — the same ephemeral, supersedable WorkflowStepProgress signal every other long-running step uses. Additionally, when EditRoomStepConfig.StreamTurns (default true), each completed turn also publishes an append-only WorkflowStepChatTurn integration event — a structural twin of WorkflowStepReasoningCaptured (sequence-numbered, every turn preserved), deliberately NOT a reuse of WorkflowStepProgress, whose "a later progress event supersedes any earlier one" semantics are wrong for a transcript where every turn matters. WorkflowStepChatTurn.IdsMentioned is server-extracted (the same regex EditRoomGroupChatManager uses for convergence checking, filtered against the offered-id set) — display/audit only, never trusted as the decision itself. The Go API/frontend relay of this new event type is a separate, follow-up pass.

Config reference​

FieldDefaultNotes
Version1Config schema version
ViewnullExtractInputRef (Previous/Step only, reused verbatim from VideoAnalyzeStepConfig) — which step's bounded {view, meta} envelope every seat and the director see. null resolves to Previous
SeatsnullIReadOnlyList<EditRoomSeat> — null/empty resolves to the three validated-live defaults (PacingEditor/StoryEditor/CraftEditor), each {Name, Persona, AgentDefinitionId}
Rounds2Full round-robin passes over every seat before the director speaks
MaxTurns8Hard ceiling (GroupChatManager.MaximumIterationCount), clamped 2..20 — NOT the framework's own default of 40
DirectorAgentDefinitionIdnullPer-agent-definition override for the director seat
TerminationSentinelOrConvergedSentinelOnly / Converged / SentinelOrConverged / FixedTurns — see the table above
MinConvergenceRounds2Consecutive rounds with an unchanged offered-id-mention set required for Converged/SentinelOrConverged to fire
MaxTurnTokens220Per-turn MaxOutputTokens, injected via RoomSeatAgent
MaxHistoryChars40000Soft cap on how much of the room transcript is rendered into the synthesis prompt (oldest turns dropped first beyond this)
Temperature0.7Sampling temperature for every editor seat's turn
DirectorTemperature0.3Sampling temperature for the director's ROOM-PARTICIPANT turns only — the standalone synthesis call uses VideoEditDirectorAgent's own AgentModelSettings default (0.3/"low") instead
ReasoningEffort"none"Injected into every seat's (and the room-participant director's) turn — measured ~3x latency reduction on this backend. Must pass ReelBoltAgentBase.ValidReasoningEfforts
RoomTimeoutSeconds1200Hard wall-clock budget for the whole group-chat run (not the later synthesis call)
PersistTranscripttrueWhether the full transcript is uploaded as ArtifactStorageKey — see above
StreamTurnstrueWhether a WorkflowStepChatTurn event is published per turn (progress reporting always happens regardless)
FallbackToSoloEditortrueWhether a failed/empty room decision falls back to one ordinary solo AgentType.VideoStoryEditor call
MaxSynthesisAttempts2Retry attempts for the standalone structured-output synthesis call

The video-derush-edit-room template​

A sixth opt-in template (AutoCreateOnProject: false), replacing video-derush-edit's middle Agent(VideoStoryEditor) step with the room: VideoAnalyze (Source: ProjectFile) → EditRoom (View: Previous) → VideoCompile (Decision: Step 2, AnalysisStepOrder: 1) → ReviewLoop (AgentType.VideoReviewAgent, looping back to step 2). The EditRoom step's own AgentDefinitionId FK is satisfied by the same AgentType.VideoTransform deterministic placeholder VideoAnalyze/ VideoCompile steps already use — the room's real seats/director are resolved independently, from EditRoomConfigJson, never from the step's own AgentDefinitionId. Deserialization-tested the same way as the other video templates (WorkflowTemplateCatalogConfigDeserializationTests.cs).


The shared room infrastructure​

The edit room's mechanics were extracted into a room-generic base the moment a second room (the graphics room, below) needed them — deliberately as ONE shared implementation, not per-room copies, since multi-agent deliberation is intended as a standing pattern in this codebase (a color-grading room is the next planned consumer). The seams:

  • RoomSeatAgent (WorkflowEngine/Agents/Rooms/RoomSeatAgent.cs, formerly EditRoomSeatAgent — renamed unchanged, it was already fully room-agnostic) + its RoomTurnResult record: the DelegatingAIAgent wrapper providing per-turn sampling-option injection, the prefix-cache-preserving persona-last message ordering, seat display names, and turn-failure containment. A new room reuses it as-is.
  • RoomGroupChatManager (WorkflowEngine/Agents/Rooms/RoomGroupChatManager.cs): the deterministic scheduler/terminator base — round-robin over N rounds then the director, the ROOM_DECIDED sentinel check, offered-id-mention convergence, the explicit ceiling check, and turn observation. A concrete room contributes ONLY its offered-id vocabulary: a thin subclass (EditRoomGroupChatManager binds [sgt]\d+, GraphicsRoomGroupChatManager binds p\d+) passing its compiled regex to the base constructor, plus a static ExtractOfferedIdMentions convenience bound to that regex.
  • IRoomStepConfig (Shared/Workflows/RoomStepConfig.cs): the config surface the shared infrastructure reads (view ref, effective seats, rounds/turn ceiling, termination knobs, temperatures, reasoning effort, timeout, transcript/stream flags, FallbackToSolo, synthesis attempts). Each room's own JSON record (EditRoomStepConfig, GraphicsRoomStepConfig) implements it on top of its unchanged JSON shape — EditRoomSeat and EditRoomTerminationMode are the room-GENERIC seat record and termination enum despite their names, kept under their original names so no persisted config or test broke when the base was extracted.
  • RoomStepExecutorBase<TDecision> (WorkflowEngine/Execution/StepExecutors/RoomStepExecutorBase.cs): the executor template — never-throws/always-valid-JSON discipline, config/view resolution, the group-chat run (agent construction, progress + WorkflowStepChatTurn events with cumulative token tallies, the TurnToken kickoff, the room timeout), transcript persistence (both the object-store artifact and the DB-persisted ChatTranscriptJson), the retried standalone synthesis call, the solo-agent fallback, offered-id filtering, and the additive "room" metadata block on output_json. A concrete room supplies: its StepType/config column/config type, offered-id extraction + mention regex, charter prompt + director turn directive, the seat/director/solo AgentTypes, decision normalize/reject/filter hooks (the edit room rejects an empty Keep list; the graphics room accepts an empty plan), and optional overrides for progress-label wording, room-turn tool scope (GetRoomTurnTools), view enrichment (PrepareViewAsync), and the empty-view outcome (BuildDecisionForEmptyView).
  • The workflow's user request reaches all three room prompts. A room does not go through StepExecutionContext.BuildAgentInput, so it does not get the --- User Request: block an ordinary StepType.Agent step's prompt ends with for free; for a while it got it nowhere, and an EditRoom step — documented as a drop-in replacement for a solo Agent(VideoStoryEditor) step — silently discarded the only channel a workflow has for creative direction. Every seeded room template then ran with RequiresUserInput: false (every template now asks for the brief), which is why nothing surfaced it until a real run asked the graphics room for two tracked inserts and the colour-grade room for a violet grade, and got zero inserts and Look: "None" with every step green. The opening message, the synthesis prompt and the solo-fallback prompt all now go through StepExecutionContext.WithUserRequest, sharing one wording with the agent-step path; the opening message is composed at the CALL SITE so a room overriding BuildOpeningMessage cannot drop it again. With no user request every prompt is byte-identical to before.

To build a new room (e.g. color grading): add a StepType + {X}RoomConfigJson jsonb column (mapped in BOTH DbContexts, migration on the WorkflowEngine context only — it owns workflow_steps), a config record implementing IRoomStepConfig, a RoomGroupChatManager subclass binding the room's id regex, a RoomStepExecutorBase<TDecision> subclass binding the hooks above, a dual-role director agent (+ seeded prompt with the verbatim-consistency test), and a template. The synthesis output schema should be an EXISTING single-agent schema whenever a solo equivalent exists, so downstream consumers need zero changes — that is the entire trick that let VideoCompileStepExecutor consume both rooms' outputs untouched.

The room convening gate — ColorGrade room only​

A room can hand a System One decision model a set of questions instead of (or alongside) just convening the group chat. The surface is IRoomStepConfig.ConveneGate (Off / Shadow / Gate, default Off everywhere) plus three virtual seams on RoomStepExecutorBase<TDecision>: BuildConveneQuestions (the wire questions a pre-room gate would ask over the bounded view), BuildConveneAgenda (the text an escalating gate would append to the opening message), and BuildShadowQuestions (the post-synthesis questions the Shadow arm records, each carrying the room's own answer as SystemTwoAnswer). All three default to null, and null is what "this room has no gate" means.

It ships for the ColorGrade room only. That is a deliberate boundary, not a partial implementation:

  • A System One choice/score/noul question can only produce a decision that decomposes into a fixed option set. ColorGradePlanOutput is exactly that — one whole-program grade in enum words — so ColorGradeRoomStepExecutor supplies the four questions AgentType.Colorist already poses (Colorist.Look / Colorist.Strength / Colorist.ShadowTone / Colorist.HighlightTone), copied verbatim from AgentStepExecutor.BuildColoristQuestions. The site strings are the solo agent's on purpose: the calibration view keys its entries by (Site, Question), so a room step's observations and a solo colorist step's land on ONE series and aggregate together. Only the recorded AgentType differs — the room's own ColorGradeDirector — which keeps a room observation attributable to the room rather than passing it off as a solo one.
  • EditRoomStepExecutor and GraphicsRoomStepExecutor override all three seams back to null, with the reason stated in the code next to the override. Neither decision decomposes into fixed options: VideoEditDecisionOutput is an id-anchored keep/drop list over whatever ids this video's analyze step happened to offer, and MotionGraphicsPlanOutput is an open, generative list of overlays whose entries carry free text and may carry a self-rendered asset key. A question set for either could only gate on a proxy for the decision, and gating on a proxy is worse than not gating: it would trade the room's real decision for a merely correlated one while looking like the room had been gated. The overrides are explicit so the absence reads as a decision rather than an oversight.
  • Their ConveneGate config member still exists and still defaults to Off, so setting it to Gate on an edit or graphics room is a no-op rather than an error — the room convenes exactly as it always has. RoomConveneGateTests asserts the emitted output is byte-identical to the same step with Off, using a strict gate mock that fails on any invocation at all.
  • Gate is live for the grade room. With ConveneGate = Gate and a Capability = Decision provider configured, the arm resolves at the dispatch point in RoomStepExecutorBase<TDecision>.ExecuteAsync — after the offered-ids gate and RecordResolvedInput, before RunRoomAsync — by calling IDecisionGate.DecideAsync with a DecisionGateRequest whose State is the bounded view alone (not the workflow's user request, not the prior-step output history), carrying SystemTwoAnswer: null and no AgentTokensUsed because nothing has run yet. It no-ops to "convene" unless the mode is Gate, an IDecisionGate was injected, and the room supplies a non-empty BuildConveneQuestions — so a null gate stays the legal, inert configuration it has always been. The gate gets its own try/catch: only OperationCanceledException on a genuinely cancelled context is rethrown, and any other failure logs and convenes.
  • A fully accepted outcome skips the room. DecisionGateOutcome.Accepted is all-or-nothing — every question cleared both AcceptAt and MinMargin — and means the group chat does not run. The step takes the room's existing RunSoloFallbackAsync, reused unchanged rather than reimplemented, so a non-convened step produces the same decision shape, the same offered-ids filtering and the same room-specific validation a degraded one does; and it advances exactly as a convened step does, through the same Success call, with room.terminationReason = "not-convened", turnCount 0 and no transcript. That is not a degrade — skipping the room is the gate working as designed — so degraded stays false and degradeReason stays null. A solo fallback that itself fails is the room's existing {FailureCodePrefix}_FAILED failure mode, never a silently emitted decision.
  • Everything else convenes the room, exactly as before. An escalation, no decision at all (outcome.HasDecision == false — no Decision provider row resolved, the gate-level deadline arriving as DecisionGateOutcome.NoDecision rather than throwing, any failure), a null IDecisionGate, and an empty question set all fall through to the ordinary group chat. On escalation the resulting agenda is appended to the opening message at the CALL SITE, after WithUserRequest and WithRetryGuidance, so a room overriding BuildOpeningMessage cannot drop it and a retried room is handed the same agenda again; null/whitespace is a strict no-op that leaves the message byte-for-byte what it was.
  • The grade room supplies the agenda through the shared formatter. ColorGradeRoomStepExecutor overrides BuildConveneAgenda to call the base's RenderConveneAgenda(BuildConveneQuestions(viewRoot), answers) — the question set is room-specific, the wording is not. Per split question, most-contested-first, it emits The fast model split on Look: Filmic 0.44 vs Warm 0.41. Resolve that disagreement first., culture-invariant and bounded by MaxConveneAgendaChars (600). That cap is a small fraction of the room's 40 000-character MaxHistoryChars deliberately: the agenda is turn zero of the group chat, so every character of it is re-sent on every turn and is rendered into the synthesis prompt alongside the bounded view under that same transcript budget — a wide question set must be able to lose its tail rather than crowd the room's own history out of it. It keeps whole sentences and then stops (never truncating an option name or a probability) and returns null when nothing qualifies, so an all-confident, empty or degenerate answer set leaves the opening message byte-for-byte what it was.
  • The Shadow arm is live too: with ConveneGate = Shadow and a Capability = Decision provider configured, the grade room records one observation per question after synthesis and changes nothing about its decision, status or transcript. No shipped configuration enables either mode — every example is commented out and Off. Off is the default everywhere and is byte-identically today's behaviour, and a room runs with no Decision provider configured at all: that is the normal local case, not an error. The default ConveneAcceptAt / ConveneMinMargin values (0.85 / 0.25) are phase-1 fallbacks, not tuned or calibrated thresholds — there is no Capability = Decision provider row and zero decision_observations rows in this environment to calibrate them against.

The graphics room​

StepType.GraphicsRoom replaces the single AgentType.MotionGraphicsPlanner planning step with a multi-agent deliberation over the SAME offered candidates a solo planner consumes — both the view.placements overlay candidates (p{n} ids, Phase 3) and the view.insertRegions tracked chroma plates (r{n} ids, Phase 5): several motion-graphics-artist seats plus a lead-artist director (AgentType.MotionGraphicsDirector) converse in a live group chat, then the director synthesizes ONE schema-validated MotionGraphicsPlanOutput — the exact schema, and the exact extended rushcut invariant (never a timestamp OR a pixel coordinate, only offered placement ids, reflection-tested by MotionGraphicsPlanOutputInvariantTests unchanged), the solo planner already produces. VideoCompileStepExecutor needs zero changes: GraphicsPlan just points at the GraphicsRoom step's StepOrder, and its deserialization skips the additive "room" metadata key exactly as Decision resolution does for the edit room.

Everything structural is the shared room infrastructure; what is specific to this room:

  • The seats (GraphicsRoomStepConfig.DefaultSeats, four by default — LayoutArtist/ TimingArtist/CopyArtist/InsertArtist) divide the actual decision space: WHERE (regions, fit scores, light/dark hints, clutter), WHEN (which candidate window lines up with what is said or shown, inEdit survival, duration words), WHAT/HOW (copy brevity, rendered-graphic vs plain text, and whether the right number of overlays is zero), and WHICH PLATES (which offered insertRegions are real screens rather than low-confidence false positives, what each should show given its aspect and surface word, and when leaving a plate empty is right). All four resolve to the same built-in AgentType.MotionGraphicsPlanner agent — personas are injected per-turn, exactly the edit room's one-agent-many-personas pattern.
  • The fourth seat is not decoration. The room originally had three, all of them arguing placements, and every other room-specific string — charter, director directive, synthesis instruction, ExtractOfferedIds, FilterToOfferedIds — was overlay-only too. Run against footage carrying two real tracked plates, with a user request explicitly asking for an insert on each, the room discussed p{n} candidates for all eight turns and synthesized an empty inserts list: two raw green rectangles in the finished film, with the room, the compile step and the execution all reporting success. Nothing owned that half of MotionGraphicsPlanOutput, so nothing raised it. The charter now describes inserts as the room's second decision, the director is told that silence on offered plates is not convergence, and the synthesis instruction names both lists.
  • Both id vocabularies, filtered separately. ExtractOfferedIds returns p{n} and r{n} together — so convergence detection covers the whole decision and a plates-only view (graphics placements off, insert tracking on) still puts the room to work rather than short-circuiting to an empty plan. FilterToOfferedIds then re-reads the two lists from the view and filters each kind against its OWN, because the combined set would wave through an overlay anchored to r0 or an insert pointing at p3.
  • AgentType.MotionGraphicsDirector is used TWICE, like VideoEditDirector: as the room-participant moderator (prose, ROOM_DECIDED), and for the standalone MotionGraphicsPlanOutput synthesis call. UNLIKE VideoEditDirector, its grant is the full sandbox+Remotion+render set minus WriteProjectFile (identical to MotionGraphicsPlanner's) — capability parity, so a room-planned overlay can still be backed by a real rendered transparent asset (RenderedAssetStorageKey) in the synthesis role. The prompt-injection tradeoff documented for MotionGraphicsPlanner (media-derived view text reaching a code-executing agent, bounded by the sandbox's containment) applies identically and is accepted for the same reasons.
  • Room turns are tool-restricted. GraphicsRoomStepExecutor.GetRoomTurnTools filters every in-room agent (seats AND the director's room instance) to the ProjectRead + WorkflowControl subset of its grant (names taken from ToolGroupCatalog, so the subset cannot drift) — a ~220-token prose turn must never reach the sandbox; only the standalone synthesis call can.
  • The view is the analyze step's placements envelope, enriched with inEdit. The template points View explicitly at the analyze step (Step 1 — Previous would resolve to the story editor's decision), and PrepareViewAsync runs the same IMotionGraphicsPlacementAnnotator a solo planner's prompt gets, read back through StepExecutionContext.OutputForPrompt. A failed annotation degrades to the plain view, never fails the step.
  • An empty plan is a VALID outcome, twice over. A view offering zero placement ids AND zero insert region ids completes immediately with an empty plan (room.terminationReason = "empty-view") instead of failing VIEW_UNRESOLVED, and a synthesized/solo plan with zero overlays (or filtered to zero by the offered-id check) completes normally — "prefer zero overlays over a cluttered edit" is the planning contract, so zero must never be treated as failure. The edit room's opposite choice (an empty Keep list degrades/fails) is the single biggest behavioral difference between the two rooms' validation hooks.
  • Failure codes: GRAPHICS_ROOM_CONFIG_INVALID / VIEW_UNRESOLVED / GRAPHICS_ROOM_FAILED / UNEXPECTED_ERROR, mirroring the edit room's. Solo fallback is one ordinary AgentType.MotionGraphicsPlanner call (FallbackToSoloPlanner, default true).

GraphicsRoomStepConfig (graphics_room_config_json, jsonb on BOTH DbContexts) carries the same knobs as EditRoomStepConfig (see Config reference) with FallbackToSoloPlanner in place of FallbackToSoloEditor, and one different default: MaxTurns is 10, not 8. The scheduler gives the seats Seats * Rounds turns and the director every turn after that, so with four seats and two rounds a ceiling of eight is consumed entirely by seat turns — the director would never speak, never moderate, and never emit ROOM_DECIDED.

The video-derush-edit-graphics-room template​

A seventh opt-in template (AutoCreateOnProject: false), replacing video-derush-edit-graphics's Agent(MotionGraphicsPlanner) step with the room: VideoAnalyze (Source: ProjectFile, emitOverlayPlacements: true) → Agent(VideoStoryEditor) → GraphicsRoom (View: Step 1) → VideoCompile (Decision: Step 2, AnalysisStepOrder: 1, enableGraphics: true, graphicsPlan: Step 3) → ReviewLoop(VideoReviewAgent) looping back to step 2. The GraphicsRoom step's own AgentDefinitionId FK is satisfied by the same AgentType.VideoTransform placeholder the other deterministic-config step types use. Deserialization-tested in WorkflowTemplateCatalogConfigDeserializationTests.cs.


Color grading​

Optional whole-program colour grading for the compiled edit, decided in enum WORDS only and resolved to concrete ffmpeg filter parameters entirely by first-party code. Two halves:

  1. The compile integration (VideoCompileStepConfig.EnableColorGrade/ColorGradePlan + ColorGradeFilterBuilder) — deterministic, opt-in, soft-failure throughout, following the EnableGraphics/GraphicsPlan and EnableMusic/MusicPlan precedent exactly.
  2. The deciders — a solo AgentType.Colorist Agent step, or the multi-agent StepType.ColorGradeRoom (below); both emit the exact same ColorGradePlanOutput shape, so the compile step consumes either interchangeably.

ColorGradePlanOutput: the words-only decision​

The rushcut invariant, extended a fourth time (after cut ids, overlay placements, and music): ColorGradePlanOutput is six plain strings — Look, Strength, ShadowTone, HighlightTone, Reason, PlanRationale — and nothing else, pinned by ColorGradePlanOutputInvariantTests (no numeric/time-bearing property; every property a string; the exact property set pinned so a future id-bearing addition fails the suite). A colorist agent is structurally incapable of emitting an RGB value, a curve point, a gamma/gain/contrast number, a percentage, or a timestamp — its entire contribution is:

WordValuesResolved to
LookNone | Warm | Cool | Filmic | Vibrant | Muted | MonoA named first-party colorbalance/eq/hue parameter set. None declines the whole grade — the compiled bytes then carry no grade filter whatsoever, tone words included
StrengthSubtle | Normal | StrongA 0.5/1.0/1.5 multiplier over the look's parameter DELTAS (Mono's hue=s=0 desaturation deliberately never scales)
ShadowToneNeutral | Lifted | DeepenedA colorlevels black-point input shift
HighlightToneNeutral | Softened | BrightenedA colorlevels white-point output/input shift

Unknown Strength/ShadowTone/HighlightTone words normalize to Normal/Neutral/Neutral (the same silent-normalize rule music applies to Intensity/Ducking); an unknown Look is never guessed at — it degrades the whole grade as unknown_look_word, because the contract's value is that every applied number traces to a recognized word.

ColorGradeFilterBuilder: the deterministic mapping​

Pure static C# (Services/Video/ColorGradeFilterBuilder.cs), exact-string-tested (ColorGradeFilterBuilderTests). Only four battle-tested core filters are ever emitted — colorbalance (colour cast), eq (contrast/saturation/gamma), hue=s=0 (Mono), colorlevels (black/white-point shaping, both tone words folded into ONE instance) — composed in that fixed order. All numbers route through FfmpegArgvFormat.Number (the R8 locale rule).

The numeric tables are deliberately INTERNAL constants, not VideoCompileStepConfig knobs: exposing raw eq/colorbalance numbers as workflow-author config would recreate, one layer up, exactly the free-numeric-parameter surface the words-only contract closes off, for no capability the curated looks don't already deliver. A future need for custom looks should add a NAMED look to the table, not a numeric pass-through. (Considered and rejected: per-look config overrides mirroring MusicBed*Db — those music knobs parameterize a level within one fixed mixing topology, while grade numbers ARE the look itself.)

Compile integration​

  • One hard failure, up front: EnableColorGrade=true with Mode=StreamCopy fails COLOR_GRADE_REQUIRES_REENCODE (a pure config error, mirroring GRAPHICS_REQUIRE_REENCODE). Everything else is soft: no_plan_configured, plan_unresolved, plan_invalid_json, unknown_look_word, look_none (the plan's own first-class no-grade decision, reported distinctly from every failure reason), grade_filters_unavailable (a probed, process-lifetime cached -filters check for colorbalance/colorlevels/hue/eq, mirroring the drawtext/amix/xfade/perspective probes), and the defensive empty_filter_chain — all degrade to "no grade applied", never a failed compile.
  • Stage ordering: grade before graphics. On the single-source select path the chain is appended to the cut stage itself (directly after setpts); on the segmented path it is its own stage directly after the concat/transition stage ([vcat]…[vgrd]). Either way it runs BEFORE any insert/overlay stage, so motion graphics always paint clean on top of graded footage — and unlike screen inserts, the grade works on BOTH encode paths (multi-source and transition-overlap compiles included), since it is a plain per-frame filter with no timeline bookkeeping to remap.
  • Reporting: the EDL and the step's output summary carry a colorGrade node (enabled/applied/look/normalized strength/shadowTone/highlightTone/filterChain/ reason) only when EnableColorGrade=true — false (the default) leaves both shapes byte-identical to the pre-grade compile path, the same load-bearing guarantee as EnableGraphics/EnableMusic/EnableInserts.

AgentType.Colorist: the solo decider​

An ordinary LLM Agent step (OutputSchemaName = "ColorGradePlanOutput"), minimal read-only tool scope (ProjectRead + WorkflowControl — identical to VideoStoryEditor/MusicSupervisor; nothing about a grade ever needs rendering, so no sandbox grant in any role). It reads the bounded VideoAnalyze view's MEASURED per-shot words — the Phase 1/Phase 4 colour temperature/tone/saturation/exposure descriptors and look groups — and picks the words above, grounding its prose Reason in shot ids (s2, s4). Reasoning disabled, temperature 0.3, same rationale as the other bounded pick-from-fixed-lists deciders.

The color grade room​

StepType.ColorGradeRoom — the third room on the shared room infrastructure: several Colorist-role seats plus a supervising colorist (AgentType.ColorGradeDirector) deliberate in a live group chat over the same bounded view a solo colorist consumes, then the director synthesizes ONE ColorGradePlanOutput OUTSIDE the chat loop. What is specific to this room:

  • The decision carries no ids at all — one whole-program grade in words. FilterToOfferedIds is a documented structural no-op (pinned by the invariant tests' property-set check), and room.droppedSpanCount is always 0. The offered s{n} SHOT ids anchor only the DELIBERATION: seats argue per-shot ("s2 reads backlit", "s0 and s4 disagree on temperature"), and shot-id mentions drive convergence detection (ColorGradeRoomGroupChatManager, pattern s\d+ — deliberately narrower than the edit room's [sgt] so a gap/segment mention never counts, and never matching p{n}/m{n}/k{n}).
  • Word validation replaces id filtering. Synthesis hard-rejects an EMPTY Look (retry with the allowlist spelled out — "no grade" must be said as the word None, never as silence); a non-empty unknown look spends one retry via CheckRetryableIssue and on the final attempt passes through to the compile step's own unknown_look_word degrade.
  • Look: "None" is a fully valid outcome (the graphics room's empty-plan grace, in this room's vocabulary), and a view offering ZERO shots completes immediately with a None plan (room.terminationReason = "empty-view") instead of failing VIEW_UNRESOLVED.
  • Seats (ColorGradeRoomStepConfig.DefaultSeats): ToneArtist (exposure/contrast, crushed or washed shots, faces), PaletteArtist (temperature/saturation words, which named look the footage wants), ContinuityArtist (one grade must suit every kept shot; when Subtle beats Strong and when None is right). All three resolve to the built-in Colorist agent — personas injected per-turn, the established one-agent-many-personas pattern.
  • No tool-restriction override is needed — unlike the graphics room, both backing agent types are minimal read-only by ToolGroupCatalog in every role, so the base default (the agent's full grant) is already the restricted set. ColorGradeDirector therefore mirrors VideoEditDirector, not MotionGraphicsDirector.
  • Failure codes: COLOR_GRADE_ROOM_CONFIG_INVALID / VIEW_UNRESOLVED / COLOR_GRADE_ROOM_FAILED / UNEXPECTED_ERROR. Solo fallback is one ordinary AgentType.Colorist call (FallbackToSoloColorist, default true).

ColorGradeRoomStepConfig (color_grade_room_config_json, jsonb on BOTH DbContexts) carries the same knobs as EditRoomStepConfig/GraphicsRoomStepConfig (see Config reference) with FallbackToSoloColorist as its fallback name. The step's AgentDefinitionId FK is satisfied by the same AgentType.VideoTransform placeholder as the other room/deterministic step types.

The video-derush-edit-grade-room template​

An eighth opt-in template (AutoCreateOnProject: false): VideoAnalyze (Source: ProjectFile — the default AnalyzeVisuals: true already emits the measured per-shot colour words, no extra flag needed) → Agent(VideoStoryEditor) → ColorGradeRoom (View: Step 1 — explicit, since Previous would resolve to the story editor's decision, not the measured-shot envelope) → VideoCompile (Decision: Step 2, AnalysisStepOrder: 1, enableColorGrade: true, colorGradePlan: Step 3) → ReviewLoop(VideoReviewAgent) looping back to step 2. Deserialization-tested in WorkflowTemplateCatalogConfigDeserializationTests.cs.

Explicitly not built (color grading)​

  • Per-shot or per-look-group grades. The v1 grade is whole-program by design: per-shot grading would need timeline-enabled filter windows remapped through every cut/transition (two disagreeing sources of timing for one filter), and per-look-group grading would mean offering the k{n} namespace to an agent — a namespace this feature deliberately keeps descriptive-only. A future per-scope phase should anchor to offered ids, not timestamps, exactly as overlays did.
  • Numeric grade knobs in config — see the ColorGradeFilterBuilder rationale above.
  • LUT files. A .cube upload path would reintroduce an opaque binary asset into the filter chain with no words-only audit trail; named first-party looks keep the EDL self-explanatory.

Sound effects​

Optional discrete sound-effect cues — whooshes, clicks, dings, transition stingers, UI sounds — mixed over the compiled edit's audio during VideoCompile's encode, each fired at the moment an offered cut-anchor item begins in the OUTPUT timeline. The same structural shape Background music established: a deterministic candidate list from VideoAnalyze, an agent that chooses among opaque offered ids plus a handful of enum words, and VideoCompileStepExecutor alone resolving those words to real ffmpeg behavior. Off by default (VideoAnalyzeStepConfig.OfferSfxClips = false, VideoCompileStepConfig.EnableSfx = false) — EnableSfx = false leaves the compile path byte-identical to the pre-SFX behavior.

Why SFX is not music with a different name​

Music and SFX share the "pick among uploaded audio/* files by opaque id" mechanics but are deliberately DIFFERENT shapes, because the use cases are inverses of each other:

Background musicSound effects
CardinalityAt most ONE track for the whole programZero or more discrete cues
DurationContinuous bed, loops/fades to fit the editShort one-shot, hard-capped by MaxSfxCueSeconds (default 4s)
PlacementProgram-wide — no moment to chooseTHE decision: an offered cut-anchor id per cue
DuckingKeyframed speech-envelope ducking under dialogueNone, deliberately — see below
No-agent pathMusicTrackProjectFileId (a fixed bed is a complete feature)None, deliberately — see below

No ducking for cues: a cue is a short, deliberately-audible accent — ducking a 300ms whoosh under dialogue would defeat its purpose, and a duck/lift envelope per cue would multiply filtergraph complexity for negative benefit. Loudness control is the Volume word plus conservative default gains (SfxSubtleDb/SfxNormalDb/SfxStrongDb, defaults -18/-12/-6 dBFS, clamped [-40, 0]), and the sound-designer prompt steers toward Subtle/Normal over dialogue.

No deterministic no-agent config path (rejected alternative): music's MusicTrackProjectFileId works because "one fixed bed under everything" is a complete feature with zero editorial judgment. The SFX equivalent would be "the same stinger at every cut" — a well-known bad pattern that would ship as an attractive footgun, while any more selective deterministic rule ("only section breaks", "only after long silences") smuggles editorial judgment into config. Cue placement — WHICH clip at WHICH moment, and whether any moment deserves one at all — is the whole decision, so the agent IS the feature here. A plan whose cues list is EMPTY is a fully valid outcome (sfx.reason = "empty_plan"), the same first-class no-op grace the graphics room's empty plan and the colorist's Look: "None" get.

Candidate discovery (VideoAnalyze)​

When OfferSfxClips = true, VideoAnalyzeStepExecutor enumerates every audio/* project file as an x{n} SFX-clip candidate (view.sfxClips, one {id, name} entry each), capped by MaxSfxClips (default 40). Project-level, not per-source, and never ffprobed at analyze time — the exact OfferMusicTracks rationale. When both OfferMusicTracks and OfferSfxClips are on, ONE ListFilesAsync call feeds both candidate lists (and one listing failure degrades both to zero candidates, never failing the step); the same file can legitimately appear under both an m{n} and an x{n} id, since nothing structural distinguishes an uploaded stinger from an uploaded bed — the sound-designer prompt tells the model to choose short one-shot clips by file name and leave bed-like names to the music layer. view.sfxClips follows the exact musicTracks budget discipline: shown as ONE ATOMIC ARRAY, not gated on VisualDetail, suppressed as a whole (after musicTracks, before insertRegions) only once detail has fully degraded, and always before any offered item is dropped. "Offered" means exactly "the whole array survived to the final view" — VideoAnalysisArtifact.OfferedSfxIds is empty whenever it was suppressed.

x{n} id isolation — and the one namespace SFX deliberately shares​

SFX-clip ids (x{n}) are their own namespace: never resolvable by VideoCompileStepExecutor.BuildIdTimeIndex (an x{n} id names a FILE, never a moment), and a Keep span naming one fails UNKNOWN_ID exactly like a placement/music/look id would. The novel bit is the cue's ANCHOR: SfxCue.AnchorId deliberately IS a cut-anchor id (s{n}/g{n}/t{n}) drawn from the same OfferedIds a Keep span may name — "this sound fires when this shot/gap/segment begins" is the whole anchoring model, reusing the one id vocabulary the model already reasons about instead of inventing per-moment SFX-placement ids. Anchor validation is against OfferedIds (offered, not merely present in the artifact), and the anchor's start is mapped through OutputTimeline.MapToOutputSec — an anchor whose moment was cut away by the edit decision drops that cue (anchor_cut_away), never relocates it.

AgentType.SoundDesigner and SfxPlanOutput​

An ordinary solo StepType.Agent step — deliberately NOT a room. The edit/graphics/grade rooms exist because those decisions have genuine multi-perspective tension across a whole program (one grade must suit every shot; overlays compete for placements and attention). An SFX cue is a small, local decision over a small candidate list, made a handful of times; multiple sound-designer personas deliberating each whoosh would multiply model calls for no perspective a single well-prompted pass lacks. If a room ever proves warranted, the shared RoomStepExecutorBase infrastructure makes it an additive follow-up, not a rewrite.

public class SfxCue
{
public string SfxId { get; set; } = ""; // must be in OfferedSfxIds; e.g. "x0"
public string AnchorId { get; set; } = ""; // must be in OfferedIds; e.g. "s2"/"g1"/"t3"
public string Timing { get; set; } = ""; // OnCut | Lead | Lag — never a number
public string Volume { get; set; } = ""; // Subtle | Normal | Strong — never a dB number
public string Reason { get; set; } = "";
}

public class SfxPlanOutput
{
public List<SfxCue> Cues { get; set; } = new(); // empty = a valid "no effects" decision
public string PlanRationale { get; set; } = "";
}

The rushcut invariant, extended a fifth time: every property is a plain string, pinned by SfxPlanOutputInvariantTests (no numeric/time-bearing property anywhere; every SfxCue property a string; both property sets pinned exactly, so a future OffsetMs-style addition fails CI). Same minimal read-only tool scope as VideoStoryEditor/MusicSupervisor — every cue plays an EXISTING uploaded clip, so there is no rendered-asset escape hatch and no sandbox grant in any role. Reasoning disabled, temperature 0.3, same rationale as the other bounded pick-from-offered-lists deciders.

Resolution: ResolveSfxAsync, soft-failure throughout​

Runs only when EnableSfx = true. Two HARD config errors up front (SFX_REQUIRES_REENCODE, SFX_REQUIRES_AUDIO_REENCODE — the exact music pair); everything else degrades:

SituationOutcome
No SfxPlan configured / unresolvable / not valid JSONNo SFX (no_plan_configured / plan_unresolved / plan_invalid_json)
Plan's cues list is emptyNo SFX, reported as the VALID empty_plan outcome, distinct from every failure
Output has no dialogue and no applied music bedThe cue is mixed over a synthesized silent base — see Program audio on silent footage. (no_base_audio used to drop every cue here, and the effects were never heard.)
This ffmpeg build's amix has no normalize optionsfx.unavailable = true, no SFX (the same cached probe music uses — normalize=0 protects the dialogue level)
Cue's SfxId not in OfferedSfxIds / AnchorId not in OfferedIdsThat cue dropped (unknown_sfx_id / unknown_anchor_id), the rest proceed
Anchor's start maps to no output moment (cut away)That cue dropped (anchor_cut_away) — never relocated
Clip not in project / not audio/* / download or probe failsThat cue dropped (sfx_not_in_project / sfx_not_audio / sfx_download_or_probe_failed)
More cues than MaxSfxCues (default 8)Excess dropped in plan order (max_cues_exceeded)

Word resolution mirrors music's exactly: unknown Timing/Volume words normalize to OnCut/Normal. Timing shifts the cue by SfxLeadMs/SfxLagMs (defaults 150ms) — applied ON THE OUTPUT TIMELINE, after the anchor mapping, so a shifted cue can never land inside a cut region its anchor's own moment survived — then clamps to [0, TotalSec]. Each cue's play window is min(clip duration, MaxSfxCueSeconds, remaining output) — a long file misused as a cue is trimmed, never allowed to become a de-facto bed — with a SfxFadeOutMs declick fade at its end (capped at half the window). Distinct clips are downloaded/probed ONCE even when several cues replay them (each cue still gets its own ffmpeg input, since each branch trims/gains/delays independently).

The mix: SfxMixFilterBuilder​

Pure static filter-fragment builder (Services/Video/SfxMixFilterBuilder.cs, exact-string-tested by SfxMixFilterBuilderTests — the MusicMixFilterBuilder role). Each cue becomes one extra -i input — always the LAST inputs, after every asset-overlay/screen-insert/music input, so every existing input-index mapping (and every filter-string test asserting one) stays untouched — and one branch:

[N:a]atrim=end={dur},asetpts=N/SR/TB,aformat=…48000…stereo,volume={gain},afade=t=out:…,adelay={ms}|{ms}[sfx{k}]

adelay (integer milliseconds, computed entirely server-side from the anchor mapping) is the one and only place a cue's timing enters the filtergraph. The mix stage layers every cue over the pre-SFX audio chain — which ends at an internal [abase] label when cues are present (the exact [aout]→[adial] label-flip convention music established), whether that base is the plain dialogue cut, the dialogue+music amix, or a music-only branch:

[abase][sfx0][sfx1]amix=inputs=3:duration=first:dropout_transition=0:normalize=0[aout]

Same three load-bearing amix options as music: normalize=0 (never quietly divide the dialogue's level), duration=first (the base pins the output length — a cue near the end can never extend the file), dropout_transition=0 (no gain re-ramp when a cue's short branch ends early, which every cue's does). The SFX mix runs AFTER the music mix and BEFORE the seam-ramp/ program-fade tail stage, so an end-of-program cue fades out with the program envelope. Unlike screen inserts (single-source-only in v1), cues work identically on BOTH encode paths — multi-source and crossfade-overlap compiles included — since adelay against the output timeline has no per-span bookkeeping to disagree with.

EDL / output JSON shape​

Present only when EnableSfx = true (byte-identical to the pre-SFX compile path otherwise):

{
"sfx": {
"enabled": true, "applied": true, "unavailable": false,
"appliedCueCount": 2,
"cues": [
{ "sfxId": "x0", "anchorId": "s4", "clipName": "whoosh.wav", "timing": "OnCut",
"volume": "Normal", "gainDb": -12, "outputStartSec": 12.4, "playDurationSec": 1.8 }
],
"dropped": [ { "reason": "anchor_cut_away", "sfxId": "x1", "anchorId": "s2" } ]
}
}

The video-derush-edit-sfx template​

A ninth opt-in template (AutoCreateOnProject: false): VideoAnalyze (Source: ProjectFile, OfferSfxClips: true) → Agent(VideoStoryEditor) → Agent(SoundDesigner) → VideoCompile (Decision: Step 2, AnalysisStepOrder: 1, EnableSfx: true, SfxPlan: Step 3) → ReviewLoop(VideoReviewAgent) looping back to step 2. Same explicit-StepOrder rationale as video-derush-edit-music (Previous relative to the compile step would resolve to the SoundDesigner step's own output, not the story editor's decision). Deserialization-tested in WorkflowTemplateCatalogConfigDeserializationTests.cs like every other template.

Explicitly not built (sound effects)​

  • A deterministic no-agent cue path — see the rejected alternative above.
  • SFX-only audio for silent outputs — now built: a silent base is synthesized, see Program audio on silent footage.
  • Per-cue ducking/sidechaining — see "no ducking for cues" above.
  • Beat-matching, auto-selected libraries, generated/synthesized effects — every cue plays an uploaded project file the model was offered by id, nothing else.

Generated clips (b-roll)​

A ShotDirector agent plans generated shots from a VideoAnalyze view; the deterministic VideoGenerate step buys one clip per shot; VideoCompile places them in the edit. This is the video-generation feature's side of the pipeline — docs/video-generation.md owns the generation side (the agent's prompt, the provider layer, the spend ledger) and this section owns what the compile side does with the result.

The template is video-derush-edit-broll (opt-in, AutoCreateOnProject: false): VideoAnalyze → Agent(VideoStoryEditor) → Agent(ShotDirector) → VideoGenerate (plan-sourced) → VideoCompile (enableGeneratedClips) → ReviewLoop(VideoReviewAgent). Every step reference in it is an explicit StepOrder, never Previous: Previous relative to the compile step would resolve to the VideoGenerate step's own output rather than the story editor's decision — the same rationale the music/SFX templates carry.

The planner's vocabulary is words and ids only​

Each planned shot is purpose (Cutaway/SeamBridge/Extend/ColdOpen/EndCard/ ScreenContent), anchor (an offered s{n}/g{n}/t{n} cut-anchor id), optional firstFrame/lastFrame (offered s{n} shot ids), camera (an enum word), duration (Short/Medium/Long), prompt (free prose) and reason. There is no number, no URL, no timestamp and no duration in seconds anywhere in it, pinned by a reflection invariant test. Every numeric thing that reaches a provider — the snapped duration, the aspect ratio, the camera clause appended to the prompt — is resolved by first-party C# in ShotPlanParameterMapper.

The prompt is the only field that leaves the building, so it goes through the shared GeneratedPromptSanitizer (control characters stripped, per-provider length cap reported rather than silently truncating) before it reaches a provider. A prompt that sanitizes to nothing, or that exceeds the cap, degrades that clip with INVALID_PROMPT — the step continues.

There is no per-purpose placement config​

Placement derives from each planned shot's own purpose, exactly as graphics and inserts derive from their plan's own items. A second config surface would only let the two disagree. The compile config gains exactly three fields, appended at the end of the positional record:

ExtractInputRef? GeneratedClips = null, // which step's resolved generated clips to place
bool EnableGeneratedClips = false, // false = byte-identical to the pre-wave compile path
int MaxGeneratedClips = 8);

EnableGeneratedClips: false — the default — is byte-identical to the pre-wave compile path: no extra ffmpeg inputs, no extra filters, no new EDL or output JSON key. That is asserted on the full argv sequence, not on the filtergraph string alone.

Clip id → storage key: the manifest, and why no path is model-visible​

A generated step produces up to MaxClips clips, and every one of them must be addressable by the compile side. WorkflowStepResult.OutputStorageKey still holds only the first succeeded clip (it drives the player in the UI), so the per-clip manifest is the real contract:

projects/{projectId}/agentFiles/video-analysis/{executionId}/step-{order}-generated-clips.json

written by VideoGenerateStepExecutor and referenced from WorkflowStepResult.ArtifactStorageKey. It carries clips[].{clipId, shotIndex, take, purpose, anchor, firstFrame, lastFrame, camera, duration, durationSeconds, status, storageKey, costUsd, costBasis, keyframeRequested, keyframeApplied, keyframeReason, failureReason, detail, skipReason} plus the step's meta, and VideoCompileStepExecutor resolves config.GeneratedClips → that step's ArtifactStorageKey → the manifest, the same way it already resolves the analysis artifact and the editorial decision. Sitting under the video-analysis/ prefix rather than outputFiles/ means the existing StepResultArtifactsController serves it with no controller change.

No storage path is model-visible. The step's OutputJson reaches later agents under AgentInputContextMode.FullWorkflow, so it must not carry an object key: the manifest's storageKey is projected out of Output and replaced by a stored boolean. That is the whole reason the manifest exists as a separate artifact instead of an extra clips[].storageKey field, and it is asserted directly — one test proves the Output contains no storageKey and none of projects/, agentFiles/ or outputFiles/, while the manifest behind ArtifactStorageKey carries the real keys.

Clip ids are v{n} (1-based plan order) when Takes is 1 and v{n}t{k} otherwise, so wave C's dailies selection has N addressable take ids to choose between. The Inline prompt path keeps phase 1's clip-{n} ids.

The placement table​

purposeWave A behaviour
CutawayReplaces picture only over the anchor's OUTPUT window — dialogue audio kept (J/L-cut)
ColdOpenPrepended as a concat segment, pairing with ProgramFadeIn
EndCardAppended as a concat segment, pairing with ProgramFadeOut
SeamBridge, Extend, ScreenContentNot applied in this wave (waves B/C) — degrade explicitly

Cutaway resolves the anchor id through BuildIdTimeIndex + OutputTimeline to its window on the output timeline, then applies the clip as an overlay on the program body's own video chain, gated by enable='between(t,start,end)' with eof_action=pass and shortest=0. The window is [anchorStart, min(anchorStart + clipDurationSeconds, anchorEnd)). Not one audio branch references a generated input — no -itsoffset, no audio shift, no shortened stream — so the anchor's dialogue plays continuously across the cutaway. An anchor whose start the edit cut away degrades that clip with anchor_not_in_output; an empty window with empty_window.

A cutaway routes the compile onto the segmented encode path even for a single source. That is the same construct a single-source crossfade compile already produces; the alternative was two cutaway implementations that can drift.

ColdOpen/EndCard are normalized to the canonical canvas and added as real concat segments (ColdOpen before segment 0, EndCard after the last), with matching-duration anullsrc synthesized when the generated clip genuinely has no audio track — probed, never assumed. The generated concat is the last content stage, immediately before the whole-piece program fade, so ProgramEnvelopeFilterBuilder resolves against the final total including both extra segments and the fades land on them. A cold open or end card that is present but unfaded because the envelope was computed before it was added would be a bug.

Everything else degrades per clip, never failing the step — the compile is never taken hostage, the same discipline graphics/music/inserts/grade/SFX already follow. A manifest that cannot be resolved at all degrades the whole feature softly (generatedClips.reason = "plan_unresolved"); the one hard config error is EnableGeneratedClips with Mode=StreamCopy (GENERATED_CLIPS_REQUIRE_REENCODE), because a picture replacement and a prepended segment both need a filtergraph. At most MaxGeneratedClips clips are applied, in manifest order; the excess degrades with max_generated_clips_exceeded. The result is reported in a generatedClips node on both the output JSON and the EDL, present only when the feature is enabled.

Egress, and what this wave cannot yet do with a keyframe​

A keyframe is not metadata about the user's footage — it is the footage. A shot naming firstFrame/lastFrame therefore needs AllowSourceMediaEgress: true, evaluated per shot by GeneratedShotEgressPolicy. A plan that asks for a frame without that consent fails as a whole step with EGRESS_REFUSED and the policy's reason: zero clips are submitted and zero ledger rows are written. A partially generated plan is a half-edit the operator cannot use, and the refusal has to be actionable (grant consent or edit the plan) rather than silently dropping one shot.

With consent granted the step proceeds. This wave cannot attach the keyframe to the request: VideoGenerationRequest has no image field, and image inputs arrive with phase 3's provider work. So a consented keyframe-requiring shot is submitted best-effort and each such clip records keyframeRequested: true, keyframeApplied: false, keyframeReason: "keyframe_transport_not_available", counted in meta.keyframesNotApplied. It is reported, never silently dropped. When a shot names a frame the config's Analysis reference is required, and the named id is checked against the analysis artifact's offered shot ids — an id that was never offered is recorded per clip as keyframe_unresolved, which is not fatal.

Budget: two estimates, one gate​

MaxSpendUsd is checked against the whole plan's estimate before anything is reserved: the pre-submit estimate sums Describe(model).PricePerSecond × snapped duration × takes across every planned shot, and exceeding the cap yields BUDGET_EXCEEDED with zero submits and zero ledger rows. That check is deliberately independent of the budget gate, so it holds even when MaxClips trims the plan to a subset that would itself have been affordable — the cap is a ceiling on the plan the author wrote, not on whatever survives the trim. The gate call that follows covers the effective (post-trim) plan, and meta.planEstimatedUsd is that effective number, equal by construction to the sum of the reservations handed to the gate. The project and global daily budgets stay owned by VideoGenerationBudgetGate, unchanged.

Not built by wave A​

  • SeamBridge, Extend and ScreenContent placements (waves B and C) — the vocabulary exists, the placements degrade explicitly.
  • Keyframe transport to the provider (phase 3), and take selection over the v{n}t{k} ids (phase 4).
  • MaxClips is still validated to 1..4, phase 1's bound, so a plan of more than four shots is trimmed rather than fully generated. Widening it is a cross-department config change (the builder mirrors the bound) and belongs with the dailies work.
  • A cold open or end card does not extend the background-music bed, the motion-graphics overlays or the tracked inserts over itself: those cover the program body. That is what keeps every overlay window and every SFX cue body-relative, so a prepended segment cannot silently shift them.

Builder UI coverage​

Every StepType the backend knows is selectable and configurable in the workflow builder — the step-type Select, the "Add Workflow Step" picker, the flowchart node map and the linear (phone) step list all cover the same eleven types, and the picker's metadata is keyed as a Record<StepType, …> so a future step type cannot compile without a UI entry.

The three room types share ONE editor and ONE canvas node (web/components/workflows/RoomStepConfigEditor.tsx, nodes/RoomNode.tsx), mirroring the way RoomStepExecutorBase<TDecision> backs all three server-side: ROOM_STEP_KINDS is the single table holding each room's config field (editRoomConfigJson / graphicsRoomConfigJson / colorGradeRoomConfigJson), its differently-named solo-fallback flag (fallbackToSoloEditor / fallbackToSoloPlanner / fallbackToSoloColorist), and its copy and accent. The editor exposes the fields that decide pipeline SHAPE (which analysis view the room deliberates over, rounds, turn ceiling, the two degrade toggles) and spread-preserves every other field, so a template-provisioned config (seats, temperatures, termination mode, timeouts) still round-trips losslessly through a save. A freshly-added or type-switched room step is written with a runnable default config immediately, since the executor hard-fails on an empty one.

VideoAnalyze's panel covers the candidate-offering switches each downstream agent needs (EmitOverlayPlacements, OfferMusicTracks, OfferSfxClips, DetectInsertRegions with its allowlisted plate colour) plus vision captioning; VideoCompile's covers every optional post-production stage (EnableGraphics, EnableInserts, EnableColorGrade, EnableMusic including the deterministic MusicTrackProjectFileId path, EnableSfx) and their plan references. Each of those toggles switches Mode back to Reencode when turned on from stream-copy, since every one of them is documented as requiring a re-encode — the backend would otherwise reject the step as a config error at execution time.

Deliberately still raw-JSON-only: per-seat persona lists, sampling temperatures, reasoning effort, room timeouts, and the numeric trim on graphics/music/SFX (per-word dB tables, overlay box geometry, lift-window shaping). Those are tuning, not pipeline shape.


Security: why ffmpeg is not in the sandbox​

ffmpeg and ffprobe run as first-party, non-AI-authored C# code inside the WorkflowEngine container. The Remotion sandbox service (/sandbox) — its allowlist, its images, its threat model — is untouched by this feature.

The sandbox's command allowlist is narrow because the code running inside it is untrusted: it executes model-authored TSX and installs model-chosen npm packages. Admitting ffmpeg to that allowlist would open one of the richest argv-injection surfaces in common Unix tooling — -i http://… (SSRF into the sandbox's bridge network, which is a normal routable Docker network, not --network none — see docs/sandbox-service.md), the concat:/subfile:/file: protocols (arbitrary in-container file read), -f lavfi with movie=, arbitrary output paths, arbitrary -map. Making that safe requires an argv validator at least as strict as simply constructing the argv yourself server-side — at which point the sandbox's containment has bought nothing, and a hardened boundary has been widened for free.

Meanwhile, the containment property the sandbox exists to provide is not needed here: this feature's ffmpeg argv is built entirely by first-party C# from a validated, typed cut list. The model's only contribution is a set of opaque ids drawn from a set the system itself issued — not one model-originated character reaches an ffmpeg argv.

Qualification (Phase 3): motion-graphics overlay TEXT is model-authored and does reach ffmpeg — but never as argv or filter-string content. It is sanitized to an allowlist, written to its own scratch file, and referenced only via drawtext's textfile= option (with expansion=none set as defense-in-depth). See Motion graphics (Phase 3) for the full discipline. The claim above — "not one model-originated character reaches an ffmpeg argv" — still holds exactly as stated for the argv/filter-string surface; it is the reason Phase 3's text still cannot inject into it.

Same boundary, the render side (B8b): the trial output policy's 720p cap and ReelBolt watermark are applied to a Remotion render by a POST-PASS in the engine — the render is pulled out of the sandbox, re-encoded once by ReactRemotionSandboxTools through the same OutputPolicyFilterBuilder chain the compile step splices in, and uploaded from the engine. The filter string is first-party C# built from two integers and a configured asset path; the model's render output never contributes an argv character here either. The sandbox still has no ffmpeg.

The sandbox's mechanics are also concretely wrong for large binary media: containers run --read-only with a 256 MB tmpfs /tmp; sandbox file I/O is base64-over-JSON (a 400 MB mp4 becomes a ~533 MB base64 string materialized in .NET memory in both directions); sandbox containers cannot reach the object store, so a video routed through the sandbox would round-trip through the engine anyway; and sandbox container lifetime is keyed to workflowExecutionId and janitor-TTL'd — the wrong lifecycle for a step that must produce a durable artifact.

Costs of running ffmpeg in-process, and how they're mitigated:

CostMitigation
CPU-heavy encoding competes with WorkflowEngine:MaxConcurrencyA singleton SemaphoreSlim in IVideoToolRunner; VideoEditing:MaxConcurrentJobs defaults to 1. The semaphore wait is cancellable and excluded from the ffmpeg timeout.
ffmpeg parses untrusted, user-uploaded media (a native attack surface)-nostdin -hide_banner -y -protocol_whitelist file on every invocation (-loglevel error for encoding; raised to info only for silence/shot detection, since those tools log their markers at ffmpeg's info level and the detector parses them from stderr); every input path is asserted to be under the per-execution scratch dir; pre-decode caps on byte size and probed duration; the process runs as the container's existing non-root $APP_UID.
Zombie processes on cancel/shutdownct.Register(() => proc.Kill(entireProcessTree: true)) plus a hard wall-clock timeout.
Container image grows (ffmpeg + its dependencies)Accepted as the cost of this design; ffmpeg is an Alpine package, not a large custom build. Measured (WS7): the built workflow-engine image is ~183 MB larger than the otherwise-identical inference image built from the same base (mcr.microsoft.com/dotnet/aspnet:9.0-alpine) — mostly ffmpeg's own codec dependency tree (libx264, libx265, libvpx, libaom, libsvtav1, vulkan loader, etc.), not the ffmpeg binary itself.

The documented phase-2 escape hatch, if this ever needs to scale independently or run with different trust boundaries, is a dedicated video-worker microservice mirroring the sandbox's per-job container model — correct long-term shape, disproportionate for a first iteration. All ffmpeg invocation sits behind a single IVideoToolRunner interface specifically so that swap is a one-class change later.


Explicitly not built​

Editing scope: no reordering of kept spans (v1 requires strictly increasing, non-overlapping spans within one clip); no arbitrary unanalyzed asset/image insertion (cutting across several pre-declared, analyzed Sources clips — including B-roll — is supported, see Multiple source clips, but inserting an image or a clip that was never fed in as a Source is not); no multicam (no automatic multi-angle sync/switching); no free-floating picture-in-picture (tracked screen inserts ARE built — a rendered scene composited into a tracked chroma-plate region, see Tracked screen inserts (Phase 5) — but an arbitrary un-tracked inset window is not); no speed ramps; no agent-requested or agent-authored transitions — only the deterministic, measurement-driven seam treatments and program fade described in Seam transitions and the program envelope.

Post scope: no colour grading / LUTs / filters / stabilization (Phase 4's D1-D3 color dimensions are measured/reported only, same as loudnorm below, never applied); no loudness normalization (loudnorm is measured and reported only, never applied); no subtitle burn-in and no SRT/VTT export (the transcript exists, so this is the most obvious phase-2 add); no speaker diarization. Background music IS built — see Background music — but it is a deterministic bed/ducking mix only: no auto-composed score, no beat-matching to cuts, no per-section music cues. Discrete sound-effect cues ARE built — see Sound effects — but only as agent-planned, id-anchored playback of uploaded clips: no generated/synthesized effects, no deterministic cue-at-every-cut path, no per-cue ducking.

Interchange: no EDL/AAF/FCPXML/OTIO export. The internal EDL JSON is an audit artifact, not an interchange format.

Timecode: no drop-frame handling, no SMPTE timecode parsing or emission, no timecode tracks. Remotion's own template renders integer 30 fps h264/mp4, and an uploaded video's rational fps (e.g. 30000/1001) is handled exactly via integer frame arithmetic without needing a timecode subsystem.

Delivery: no streaming/HLS packaging, no proxy/preview transcodes, no thumbnail sprites, no waveform PNGs.

Approval: no suspend/resume human-approval gate mid-execution — "approval by composition" (run analysis + decision, inspect, then run compile separately against the saved artifact) instead. A real suspend/resume gate is the single highest-value phase-2 item.

Platform: no changes to /sandbox whatsoever — not the Go source, not the allowlist, not the images, not the compose entry. No new microservice in v1. No async/polling execution model beyond what the workflow engine already has. No ProjectFile backfill for historical render outputs.

Limits: MaxDurationSeconds defaults to 1800 (30 min), MaxInputBytes to 2 GB — revisit if real inputs are much longer or shorter. No GPU encoder path (e.g. h264_nvenc) unless the deployment host is confirmed to have one.